Skip to content

Eve Fleisig

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

12

Venues

3

Active years

2023–2026

Best venue rank

A*

Where they publish

Papers

12 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLAI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Zhu Irene Ying, Tianyi Zhou, Jordan Lee Boyd-Graber
2025ACLGRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration.Yoo Yeon Sung, Eve Fleisig, Yu Hou, Ishan Upadhyay, Jordan Lee Boyd-Graber
2025NAACLIs your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness.Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, Jordan Lee Boyd-Graber
2024EMNLPLinguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination.Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, Dan Klein
2024EMNLPAccurate and Data-Efficient Toxicity Prediction when Annotators Disagree.Harbani Jaggi, Kashyap Coimbatore Murali, Eve Fleisig, Erdem Biyik
2024NAACLThe Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.Eve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat
2024NAACLFirst Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models.Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez
2024NAACLGhostbuster: Detecting Text Ghostwritten by Large Language Models.Vivek Verma, Eve Fleisig, Nicholas Tomlin, Dan Klein
2023ACLFairPrism: Evaluating Fairness-Related Harms in Text Generation.Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daum III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna M. Wallach
2023EMNLPWhen the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks.Eve Fleisig, Rediet Abebe, Dan Klein
2023EMNLPIncorporating Worker Perspectives into MTurk Annotation Practices for NLP.Olivia Huang, Eve Fleisig, Dan Klein
2023EMNLPCentering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection.Vyoma Raman, Eve Fleisig, Dan Klein