Skip to content

RL with KL penalties is better viewed as Bayesian inference.

Tomasz Korbak, Ethan Perez, Christopher L. Buckley

VenueA*EMNLP
Year2022
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.