Skip to content

DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF.

Ziyuan Gao, Di Liang, Xianjie Wu, Philippe Morel, Minlong Peng

VenueA*AAAI
Year2026
ProceedingsAAAI

Browse the full AAAI paper archive.