Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug Errors.
Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
Browse the full EMNLP paper archive.
Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
Browse the full EMNLP paper archive.