Skip to content

A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models.

Hayeon Lee, Rui Hou, Jongpil Kim, Davis Liang, Sung Ju Hwang, Alexander Min

VenueA*ACL
Year2023
ProceedingsACL (Findings)

Browse the full ACL paper archive.