Skip to content

Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning.

Wei Yang, Jesse Thomason

VenueA*AAAI
Year2026
ProceedingsAAAI

Browse the full AAAI paper archive.