| 2025 | ACL | BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs. | Jesse Woo, Fateme Hashemi Chaleshtori, Ana Marasovic, Kenneth Marino |
| 2025 | EMNLP | What Has Been Lost with Synthetic Evaluation? | Alexander Gill, Abhilasha Ravichander, Ana Marasovic |
| 2025 | EMNLP | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps. | Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov |
| 2024 | EMNLP | On Evaluating Explanation Utility for Human-AI Decision Making in NLP. | Fateme Hashemi Chaleshtori, Atreya Ghosal, Alexander Gill, Purbid Bambroo, Ana Marasovic |
| 2024 | NAACL | Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness. | Ashim Gupta, Rishanth Rajendhran, Nathan Stringham, Vivek Srikumar, Ana Marasovic |
| 2023 | ACL | Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest. | Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi |
| 2022 | EMNLP | On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization. | Shruti Palaskar, Akshita Bhagia, Yonatan Bisk, Florian Metze, Alan W. Black, Ana Marasovic |
| 2022 | EMNLP | CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation. | Abhilasha Ravichander, Matt Gardner, Ana Marasovic |
| 2022 | EMNLP | Does Self-Rationalization Improve Robustness to Spurious Correlations? | Alexis Ross, Matthew E. Peters, Ana Marasovic |
| 2022 | NAACL | Few-Shot Self-Rationalization with Natural Language Prompts. | Ana Marasovic, Iz Beltagy, Doug Downey, Matthew E. Peters |
| 2021 | ACL | Promoting Graph Awareness in Linearized Graph-to-Text Generation. | Alexander Miserlis Hoyle, Ana Marasovic, Noah A. Smith |
| 2021 | ACL | Explaining NLP Models via Minimal Contrastive Editing (MiCE). | Alexis Ross, Ana Marasovic, Matthew E. Peters |
| 2021 | ACL | Effective Attention Sheds Light On Interpretability. | Kaiser Sun, Ana Marasovic |
| 2021 | EMNLP | Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. | Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner |
| 2021 | EMNLP | Measuring Association Between Labels and Free-Text Rationales. | Sarah Wiegreffe, Ana Marasovic, Noah A. Smith |
| 2020 | ACL | Don't Stop Pretraining: Adapt Language Models to Domains and Tasks. | Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith |
| 2020 | EMNLP | Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs. | Ana Marasovic, Chandra Bhagavatula, Jae Sung Park, Ronan Le Bras, Noah A. Smith, Yejin Choi |
| 2020 | EMNLP | Easy, Reproducible and Quality-Controlled Data Collection with CROWDAQ. | Qiang Ning, Hao Wu, Pradeep Dasigi, Dheeru Dua, Matt Gardner, Robert L. Logan IV, Ana Marasovic, Zhen Nie |
| 2019 | EMNLP | Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning. | Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, Matt Gardner |
| 2018 | NAACL | SRL4ORL: Improving Opinion Role Labeling Using Multi-Task Learning with Semantic Role Labeling. | Ana Marasovic, Anette Frank |
| 2017 | EMNLP | A Mention-Ranking Model for Abstract Anaphora Resolution. | Ana Marasovic, Leo Born, Juri Opitz, Anette Frank |