Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and Mixture.
Tianhao Fu, Xinxin Xu, Weichen Xu, Jue Chen, Ruilong Ren, Bowen Deng, Xinyu Zhao, Jian Cao, Xixin Cao
Browse the full AAAI paper archive.