Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.
Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, Jingren Zhou
Browse the full ICLR paper archive.
Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, Jingren Zhou
Browse the full ICLR paper archive.