DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF.
Ziyuan Gao, Di Liang, Xianjie Wu, Philippe Morel, Minlong Peng
Browse the full AAAI paper archive.
Ziyuan Gao, Di Liang, Xianjie Wu, Philippe Morel, Minlong Peng
Browse the full AAAI paper archive.