OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models.
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, Yulan He
Browse the full ACL paper archive.
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, Yulan He
Browse the full ACL paper archive.