Skip to content

Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning.

Tianduo Wang, Shichen Li, Wei Lu

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.