R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling.
Aijia Cheng, Kailong Wang, Ling Shi, Yongxin Zhao
Browse the full ACL paper archive.
Aijia Cheng, Kailong Wang, Ling Shi, Yongxin Zhao
Browse the full ACL paper archive.