Skip to content

Linear attention is (maybe) all you need (to understand Transformer optimization).

Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, Suvrit Sra

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.