| 2025 | QCoder Benchmark: Bridging Language Generation and Quantum Hardware through Simulator-Based Feedback. | Taku Mikuriya, Tatsuya Ishigaki, Masayuki Kawarada, Shunya Minami, Tadashi Kadowaki, Yohichi Suzuki, Soshun Naito, Shunya Takada, Takumi Kato, Tamotsu Basseda, Reo Yamada, Hiroya Takamura |
| 2025 | Can LLMs Help Encoder Models Maintain Both High Accuracy and Consistency in Temporal Relation Classification? | Adiel Meir, Kfir Bar |
| 2025 | Cognitive Flow: An LLM-Automated Framework for Quantifying Reasoning Distillation. | Jos Matos, Catarina Silva, Hugo Gonalo Oliveira |
| 2025 | Exploring the Power of Large Language Models for Vietnamese Implitcit Sentiment Analysis. | Huy Gia Luu, Dang Van Thin |
| 2025 | ViNumFCR: A Novel Vietnamese Benchmark for Numerical Reasoning Fact Checking on Social Media News. | Nhi Ngoc Phuong Luong, Anh Thi Lan Le, Tin Van Huynh, Kiet Van Nguyen, Ngan Nguyen |
| 2025 | Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues. | Dongxu Lu, Johan Jeuring, Albert Gatt |
| 2025 | Who's Laughing Now? An Overview of Computational Humour Generation and Explanation. | Tyler Loakman, William Thorne, Chenghua Lin |
| 2025 | Annotating Hallucinations in Question-Answering using Rewriting. | Xu Liu, Guanyi Chen, Kees van Deemter, Tingting He |
| 2025 | Counterfactual Simulatability of LLM Explanations for Generation Tasks. | Marvin Limpijankit, Yanda Chen, Melanie Subbiah, Nicholas Deas, Kathleen McKeown |
| 2025 | Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation. | Tomasz Limisiewicz, David Marecek, Toms Musil |
| 2025 | When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition. | Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek, Ehud Reiter |
| 2025 | Restaurant Menu Categorization at Scale: LLM-Guided Hybrid Clustering. | Seemab Latif, Ashar Mehmood, Selim Turki, Huma Ameer, Ivan Gorban, Faysal Fateh |
| 2025 | OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs. | Ivan Kartc, Mateusz Lango, Ondrej Dusek |
| 2025 | Generating Impact and Critique Explanations of Predictions made by a Goal Recognizer. | Jair da Silva Ferreira Junior, Ingrid Zukerman, Enes Makalic, Ccile Paris, Mor Vered |
| 2025 | Surprisal reveals diversity gaps in image captioning and different scorers change the story. | Nikolai Ilinykh, Simon Dobnik |
| 2025 | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics. | Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
| 2025 | Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs. | Akio Hayakawa, Stefan Bott, Horacio Saggion |
| 2025 | Natural Language Translation of Formal Proofs through Informalization of Proof Steps and Recursive Summarization along Proof Structure. | Seiji Hattori, Takuya Matsuzaki, Makoto Fujiwara |
| 2025 | Human ratings of LLM response generation in pair-programming dialogue. | Cecilia Domingo, Paul Piwek, Svetlana Stoyanchev, Michel Wermelinger, Kaustubh Adhikari, Rama Sanand Doddipatla |
| 2025 | Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models. | Cong-Thanh Do, Rama Sanand Doddipatla, Kate M. Knill |
| 2025 | From Prototypical to Relational: How LLMs Navigate Complex Analogies. | Mayukh Das, Wolf-Tilo Balke |
| 2025 | References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation. | Silvia Casola, Yang Janet Liu, Siyao Peng, Oliver Kraus, Albert Gatt, Barbara Plank |
| 2025 | Incorporating Formulaicness in the Automatic Evaluation of Naturalness: A Case Study in Logic-to-Text Generation. | Eduardo Cal, Guanyi Chen, Elias Stengel-Eskin, Albert Gatt, Kees van Deemter |
| 2025 | Statistical Multicriteria Evaluation of LLM-Generated Text. | Esteban Garces Arias, Hannah Blocher, Julian Rodemann, Matthias Aenmacher, Christoph Jansen |
| 2025 | Evaluating LLMs' Ability to Understand Numerical Time Series for Text Generation. | Mizuki Arai, Tatsuya Ishigaki, Masayuki Kawarada, Yusuke Miyao, Hiroya Takamura, Ichiro Kobayashi |