Skip to content

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, Olivier Bachem

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.