Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty
Browse the full ACL paper archive.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty
Browse the full ACL paper archive.