Skip to content

Gated Linear Attention Transformers with Hardware-Efficient Training.

Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.