Skip to content

ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback.

Ju-Seung Byun, Jiyun Chun, Jihyung Kil, Andrew Perrault

VenueA*EMNLP
Year2024
ProceedingsEMNLP

Browse the full EMNLP paper archive.