How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation Under the One-Time-Pad-Based Framework.
Zi Liang, Liantong Yu, Shiyu Zhang, Qingqing Ye, Haibo Hu
Browse the full AAAI paper archive.
Zi Liang, Liantong Yu, Shiyu Zhang, Qingqing Ye, Haibo Hu
Browse the full AAAI paper archive.