Efficient LLMs with AMP: Attention Heads and MLP Pruning.
Leandro Giusti Mugnaini, Bruno Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Edson Bollis, Lucas F. A. O. Pellicer, Anna Helena Reali Costa, Artur Jordo
Browse the full IJCNN paper archive.