| 2026 | ACL | Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine. | Sebastian Antony Joseph, Lily Chen, Barry Wei, Michael Mackert, Iain James Marshall, Paul Pu Liang, Ramez Kouzy, Byron C. Wallace, Junyi Jessy Li |
| 2026 | ACL | Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. | Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li |
| 2025 | ACL | Who Taught You That? Tracing Teachers in Model Distillation. | Somin Wadhwa, Chantal Shaib, Silvio Amir, Byron C. Wallace |
| 2025 | EMNLP | Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare. | Hiba Ahsan, Arnab Sen Sharma, Silvio Amir, David Bau, Byron C. Wallace |
| 2025 | ICLR | NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals. | Jaden Fried Fiotto-Kaufman, Alexander Russell Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell, Byron C. Wallace, David Bau |
| 2024 | ACL | FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence. | Sebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Gke, Monika Coers, Wei Xu, Byron C. Wallace, Junyi Jessy Li |
| 2024 | ACL | InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification. | Jan Trienes, Sebastian Joseph, Jrg Schltterer, Christin Seifert, Kyle Lo, Wei Xu, Byron C. Wallace, Junyi Jessy Li |
| 2024 | EACL | Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains. | Sanjana Ramprasad, Kundan Krishna, Zachary C. Lipton, Byron C. Wallace |
| 2024 | EACL | Leveraging ChatGPT in Pharmacovigilance Event Extraction: An Empirical Study. | Zhaoyue Sun, Gabriele Pergola, Byron C. Wallace, Yulan He |
| 2024 | EMNLP | Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs. | Sheridan Feucht, David Atkinson, Byron C. Wallace, David Bau |
| 2024 | EMNLP | Detection and Measurement of Syntactic Templates in Generated Text. | Chantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. Wallace |
| 2024 | EMNLP | Investigating Mysteries of CoT-Augmented Distillation. | Somin Wadhwa, Silvio Amir, Byron C. Wallace |
| 2024 | EMNLP | Learning from Natural Language Explanations for Generalizable Entity Matching. | Somin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace, Luyang Kong |
| 2024 | ICLR | Evaluating the Zero-shot Robustness of Instruction-tuned Language Models. | Jiuding Sun, Chantal Shaib, Byron C. Wallace |
| 2024 | ICLR | Function Vectors in Large Language Models. | Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, David Bau |
| 2024 | NAACL | Towards Reducing Diagnostic Errors with Interpretable Risk Prediction. | Denis Jered McInerney, William Dickinson, Lucy C. Flynn, Andrea Young, Geoffrey S. Young, Jan-Willem van de Meent, Byron C. Wallace |
| 2024 | NAACL | On-the-fly Definition Augmentation of LLMs for Biomedical NER. | Monica Munnangi, Sergey Feldman, Byron C. Wallace, Silvio Amir, Tom Hope, Aakanksha Naik |
| 2023 | ACL | Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT-3 (with Varying Success). | Chantal Shaib, Millicent L. Li, Sebastian Joseph, Iain James Marshall, Junyi Jessy Li, Byron C. Wallace |
| 2023 | ACL | Revisiting Relation Extraction in the era of Large Language Models. | Somin Wadhwa, Silvio Amir, Byron C. Wallace |
| 2023 | ACL | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations. | Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong, Bailey Kuehl, Erin Bransom, Byron C. Wallace |
| 2023 | CoNLL | Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. | Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, David Bau |
| 2023 | EACL | NapSS: Paragraph-level Medical Text Simplification via Narrative Prompting and Sentence-matching Summarization. | Junru Lu, Jiazheng Li, Byron C. Wallace, Yulan He, Gabriele Pergola |
| 2023 | EACL | Automatically Summarizing Evidence from Clinical Trials: A Prototype Highlighting Current Challenges. | Sanjana Ramprasad, Denis Jered McInerney, Iain James Marshall, Byron C. Wallace |
| 2023 | EACL | RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media. | Somin Wadhwa, Vivek Khetan, Silvio Amir, Byron C. Wallace |
| 2023 | EACL | How Many and Which Training Points Would Need to be Removed to Flip this Prediction? | Jinghan Yang, Sarthak Jain, Byron C. Wallace |
| 2023 | EMNLP | Multilingual Simplification of Medical Texts. | Sebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan, Wei Xu, Byron C. Wallace, Junyi Jessy Li |
| 2023 | EMNLP | USB: A Unified Summarization Benchmark Across Tasks and Domains. | Kundan Krishna, Prakhar Gupta, Sanjana Ramprasad, Byron C. Wallace, Jeffrey P. Bigham, Zachary C. Lipton |
| 2023 | EMNLP | CHiLL: Zero-shot Custom Interpretable Feature Extraction from Clinical Notes with Large Language Models. | Denis Jered McInerney, Geoffrey S. Young, Jan-Willem van de Meent, Byron C. Wallace |
| 2023 | EMNLP | Appraising the Potential Uses and Harms of LLMs for Medical Systematic Reviews. | Hye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. Wallace |
| 2023 | IVA | Accomodating User Expressivity while Maintaining Safety for a Virtual Alcohol Misuse Counselor. | Stefan Olafsson, Paola Pedrelli, Byron C. Wallace, Timothy W. Bickmore |
| 2022 | ACL | Evaluating Factuality in Text Simplification. | Ashwin Devaraj, William Sheffield, Byron C. Wallace, Junyi Jessy Li |
| 2022 | ACL | Combining Feature and Instance Attribution to Detect Artifacts. | Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, Byron C. Wallace |
| 2022 | COLING | Overview of MSLR2022: A Shared Task on Multi-document Summarization for Literature Reviews. | Lucy Lu Wang, Jay DeYoung, Byron C. Wallace |
| 2022 | EMNLP | Influence Functions for Sequence Tagging Models. | Sarthak Jain, Varun Manjunatha, Byron C. Wallace, Ani Nenkova |
| 2022 | EMNLP | That's the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical Data. | Denis Jered McInerney, Geoffrey S. Young, Jan-Willem van de Meent, Byron C. Wallace |
| 2022 | EMNLP | PHEE: A Dataset for Pharmacovigilance Event Extraction from Text. | Zhaoyue Sun, Jiazheng Li, Gabriele Pergola, Byron C. Wallace, Bino John, Nigel Greene, Joseph Kim, Yulan He |
| 2022 | IJCNLP | Self-Repetition in Abstractive Neural Summarizers. | Nikita Salkar, Thomas A. Trikalinos, Byron C. Wallace, Ani Nenkova |
| 2021 | ACL | Biomedical Interpretable Entity Representations. | Diego Garcia-Olano, Yasumasa Onoe, Ioana Baldini, Joydeep Ghosh, Byron C. Wallace, Kush R. Varshney |
| 2021 | AMIA | Identifying Communication Behavior Indicators in Secure Messages. | Dezon K. Finch, Lina Bouayad, Timothy P. Hogan, Sarah L. Cutrona, Byron C. Wallace, Stephen L. Luther, Bridget Smith, Stephanie L. Shimada |
| 2021 | AMIA | Applying State of the Art Language Models to Enable Better Clinical Natural Language Processing. | Bryan D. Steitz, Emily Alsentzer, Hoo Chang Shin, Byron C. Wallace, Adam Wright |
| 2021 | EMNLP | Unsupervised Data Augmentation with Naive Augmentation and without Unlabeled Data. | David Lowell, Brian E. Howard, Zachary C. Lipton, Byron C. Wallace |
| 2021 | EMNLP | Disentangling Representations of Text by Masking Transformers. | Xiongyi Zhang, Jan-Willem van de Meent, Byron C. Wallace |
| 2021 | NAACL | On the Impact of Random Seeds on the Fairness of Clinical Classifiers. | Silvio Amir, Jan-Willem van de Meent, Byron C. Wallace |
| 2021 | NAACL | Paragraph-level Simplification of Medical Texts. | Ashwin Devaraj, Iain James Marshall, Byron C. Wallace, Junyi Jessy Li |
| 2021 | NAACL | Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? | Eric P. Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, Byron C. Wallace |
| 2021 | NAACL | An Empirical Comparison of Instance Attribution Methods for NLP. | Pouya Pezeshkpour, Sarthak Jain, Byron C. Wallace, Sameer Singh |
| 2020 | ACL | ERASER: A Benchmark to Evaluate Rationalized NLP Models. | Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace |
| 2020 | ACL | Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions. | Xiaochuang Han, Byron C. Wallace, Yulia Tsvetkov |
| 2020 | ACL | Learning to Faithfully Rationalize by Construction. | Sarthak Jain, Sarah Wiegreffe, Yuval Pinter, Byron C. Wallace |
| 2020 | ACL | Trialstreamer: Mapping and Browsing Medical Evidence in Real-Time. | Benjamin E. Nye, Ani Nenkova, Iain James Marshall, Byron C. Wallace |
| 2019 | AISTATS | Structured Neural Topic Models for Reviews. | Babak Esmaeili, Hongyi Huang, Byron C. Wallace, Jan-Willem van de Meent |
| 2019 | EMNLP | Practical Obstacles to Deploying Active Learning. | David Lowell, Zachary C. Lipton, Byron C. Wallace |
| 2019 | IJCAI | What Does the Evidence Say? Models to Help Make Sense of the Biomedical Literature. | Byron C. Wallace |
| 2019 | IUI | Explainable modeling of annotations in crowdsourcing. | An T. Nguyen, Matthew Lease, Byron C. Wallace |
| 2019 | IUI | Mash: software tools for developing interactive and transparent machine learning systems. | An Thanh Nguyen, Matt Lease, Byron C. Wallace |
| 2019 | NAACL | Attention is not Explanation. | Sarthak Jain, Byron C. Wallace |
| 2019 | NAACL | Inferring Which Medical Treatments Work from Reports of Clinical Trials. | Eric P. Lehman, Jay DeYoung, Regina Barzilay, Byron C. Wallace |
| 2019 | NAACL | Predicting Annotation Difficulty to Improve Task Routing and Model Performance for Biomedical Information Extraction. | Yinfei Yang, Oshin Agarwal, Chris Tar, Byron C. Wallace, Ani Nenkova |
| 2018 | AAAI | An Interpretable Joint Graphical Model for Fact-Checking From Crowds. | An T. Nguyen, Aditya Kharosekar, Matthew Lease, Byron C. Wallace |
| 2018 | ACL | A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature. | Benjamin E. Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain James Marshall, Ani Nenkova, Byron C. Wallace |
| 2018 | EMNLP | Learning Disentangled Representations of Texts with Application to Biomedical Abstracts. | Sarthak Jain, Edward Banner, Jan-Willem van de Meent, Iain James Marshall, Byron C. Wallace |
| 2018 | EMNLP | Structured Multi-Label Biomedical Text Tagging via Attentive Neural Tree Decoding. | Gaurav Singh, James Thomas, Iain James Marshall, John Shawe-Taylor, Byron C. Wallace |
| 2018 | NAACL | Syntactic Patterns Improve Information Extraction for Medical Search. | Roma Patel, Yinfei Yang, Iain James Marshall, Ani Nenkova, Byron C. Wallace |
| 2018 | SIGIR | Automating Biomedical Evidence Synthesis: Recent Work and Directions Forward. | Byron C. Wallace |
| 2018 | UIST | Believe it or not: Designing a Human-AI Partnership for Mixed-Initiative Fact-Checking. | An T. Nguyen, Aditya Kharosekar, Saumyaa Krishnan, Siddhesh Krishnan, Elizabeth Tate, Byron C. Wallace, Matthew Lease |
| 2017 | AAAI | Active Discriminative Text Representation Learning. | Ye Zhang, Matthew Lease, Byron C. Wallace |
| 2017 | ACL | Automating Biomedical Evidence Synthesis: RobotReviewer. | Iain James Marshall, Jol Kuiper, Edward Banner, Byron C. Wallace |
| 2017 | ACL | Aggregating and Predicting Sequence Labels from Crowd Annotations. | An Thanh Nguyen, Byron C. Wallace, Junyi Jessy Li, Ani Nenkova, Matthew Lease |
| 2017 | ACL | Exploiting Domain Knowledge via Grouped Weight Sharing with Application to Text Categorization. | Ye Zhang, Matthew Lease, Byron C. Wallace |
| 2017 | AMIA | Detecting Twitter posts with Adverse Drug Reactions using Convolutional Neural Networks. | Sarthak Jain, Xun Peng, Byron C. Wallace |
| 2017 | CIKM | A Neural Candidate-Selector Architecture for Automatic Structured Clinical Text Annotation. | Gaurav Singh, Iain James Marshall, James Thomas, John Shawe-Taylor, Byron C. Wallace |
| 2017 | IJCNLP | A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification. | Ye Zhang, Byron C. Wallace |
| 2016 | CoNLL | Modelling Context with User Embeddings for Sarcasm Detection in Social Media. | Silvio Amir, Byron C. Wallace, Hao Lyu, Paula Carvalho, Mrio J. Silva |
| 2016 | EMNLP | Rationale-Augmented Convolutional Neural Networks for Text Classification. | Ye Zhang, Iain James Marshall, Byron C. Wallace |
| 2016 | HCOMP | Probabilistic Modeling for Crowdsourcing Partially-Subjective Ratings. | An Thanh Nguyen, Matthew Halpern, Byron C. Wallace, Matthew Lease |
| 2016 | NAACL | MGNC-CNN: A Simple Approach to Exploiting Multiple Word Embeddings for Sentence Classification. | Ye Zhang, Stephen Roller, Byron C. Wallace |
| 2016 | UAI | A Correlated Worker Model for Grouped, Imbalanced and Multitask Data. | An T. Nguyen, Byron C. Wallace, Matthew Lease |
| 2015 | AAAI | Graph-Sparse LDA: A Topic Model with Structured Sparsity. | Finale Doshi-Velez, Byron C. Wallace, Ryan P. Adams |
| 2015 | AAAI | What Predicts Media Coverage of Health Science Articles? | Byron C. Wallace, Michael J. Paul, Nomie Elhadad |
| 2015 | ACL | Sparse, Contextually Informed Models for Irony Detection: Exploiting User Communities, Entities and Sentiment. | Byron C. Wallace, Do Kook Choe, Eugene Charniak |
| 2015 | AMIA | Improving Retrieval of PubMed Articles Using the TopicalMeSH Representation. | Zhiguo Yu, Elmer V. Bernstam, Trevor Cohen, Byron C. Wallace, Todd R. Johnson |
| 2015 | HCOMP | Combining Crowd and Expert Labels Using Decision Theoretic Active Learning. | An Thanh Nguyen, Byron C. Wallace, Matthew Lease |
| 2014 | AAAI | Preface. | Finale Doshi-Velez, David C. Kale, Byron C. Wallace, Jenna Wiens |
| 2014 | AAAI | Discovering Better AAAI Keywords via Clustering with Community-Sourced Constraints. | Kelly Moran, Byron C. Wallace, Carla E. Brodley |
| 2014 | AAAI | Organizers. | Byron C. Wallace |
| 2014 | AAAI | Identifying Differences in Physician Communication Styles with a Log-Linear Transition Component Model. | Byron C. Wallace, Issa J. Dahabreh, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
| 2014 | ACL | Humans Require Context to Infer Ironic Intent (so Computers Probably do, too). | Byron C. Wallace, Do Kook Choe, Laura Kertz, Eugene Charniak |
| 2014 | CogSci | Can Cognitive Scientists Help Computers Recognize Irony? | Byron C. Wallace, Laura Kertz |
| 2013 | EMNLP | A Generative Joint, Additive, Sequential Model of Topics and Speech Acts in Patient-Doctor Communication. | Byron C. Wallace, Thomas A. Trikalinos, Michael Barton Laws, Ira B. Wilson, Eugene Charniak |
| 2012 | ICDM | Class Probability Estimates are Unreliable for Imbalanced Data (and How to Fix Them). | Byron C. Wallace, Issa J. Dahabreh |
| 2012 | NAACL | Multiple Narrative Disentanglement: Unraveling Infinite Jest. | Byron C. Wallace |
| 2011 | ICDM | Class Imbalance, Redux. | Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
| 2011 | ICML | The Constrained Weight Space SVM: Learning with Ranked Features. | Kevin Small, Byron C. Wallace, Carla E. Brodley, Thomas A. Trikalinos |
| 2011 | SDM | Who Should Label What? Instance Allocation in Multiple Expert Active Learning. | Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
| 2010 | KDD | Active learning for biomedical citation screening. | Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |