| 2026 | ACL | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models. | Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka |
| 2025 | EMNLP | Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict. | Pouya Pezeshkpour, Moin Aminnaseri, Estevam Hruschka |
| 2025 | NAACL | From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization. | Catarina G. Belm, Pouya Pezeshkpour, Hayate Iso, Seiji Maekawa, Nikita Bhutani, Estevam Hruschka |
| 2025 | NAACL | LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs. | Arash Gholami Davoodi, Seyed Pouyan Mousavi Davoudi, Pouya Pezeshkpour |
| 2025 | NAACL | Evaluating Bias in LLMs for Job-Resume Matching: Gender, Race, and Education. | Hayate Iso, Pouya Pezeshkpour, Nikita Bhutani, Estevam Hruschka |
| 2025 | NAACL | Multi-Conditional Ranking with Large Language Models. | Pouya Pezeshkpour, Estevam Hruschka |
| 2024 | EACL | Less is More for Long Document Summary Evaluation by LLMs. | Yunshu Wu, Hayate Iso, Pouya Pezeshkpour, Nikita Bhutani, Estevam Hruschka |
| 2024 | NAACL | Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. | Pouya Pezeshkpour, Estevam Hruschka |
| 2023 | ICMLA | Measuring and Modifying Factual Knowledge in Large Language Models. | Pouya Pezeshkpour |
| 2022 | ACL | Combining Feature and Instance Attribution to Detect Artifacts. | Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, Byron C. Wallace |
| 2021 | NAACL | An Empirical Comparison of Instance Attribution Methods for NLP. | Pouya Pezeshkpour, Sarthak Jain, Byron C. Wallace, Sameer Singh |
| 2019 | NAACL | Investigating Robustness and Interpretability of Link Prediction via Adversarial Modifications. | Pouya Pezeshkpour, Yifan Tian, Sameer Singh |
| 2018 | EMNLP | Embedding Multimodal Relational Data for Knowledge Base Completion. | Pouya Pezeshkpour, Liyan Chen, Sameer Singh |