Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences.
Zicheng Liu, Siyuan Li, Li Wang, Zedong Wang, Yunfan Liu, Stan Z. Li
Browse the full ICML paper archive.
Zicheng Liu, Siyuan Li, Li Wang, Zedong Wang, Yunfan Liu, Stan Z. Li
Browse the full ICML paper archive.