Skip to content

Katherine Lee

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

15

Venues

8

Active years

2004–2025

Best venue rank

A*

Where they publish

Papers

15 indexed papers, newest first.

YearVenueTitleAuthors
2025ACLPrivacy Ripple Effects from Adding or Removing Personal Information in Language Model Training.Jaydeep Borkar, Matthew Jagielski, Katherine Lee, Niloofar Mireshghallah, David A. Smith, Christopher A. Choquette-Choo
2025ICLRScalable Extraction of Training Data from Aligned, Production Language Models.Milad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Florian Tramr, Katherine Lee
2025ICLRRecite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon.USVSN Sai Prashanth, Alvin Deng, Kyle O'Brien, Jyothir S. V, Mohammad Aflah Khan, Jaydeep Borkar, Christopher A. Choquette-Choo, Jacob Ray Fuehne, Stella Biderman, Tracy Ke, Katherine Lee, Naomi Saphra
2025ICMLExploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards.Yangsibo Huang, Milad Nasr, Anastasios Nikolas Angelopoulos, Nicholas Carlini, Wei-Lin Chiang, Christopher A. Choquette-Choo, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Ken Liu, Ion Stoica, Florian Tramr, Chiyuan Zhang
2025NAACLMeasuring memorization in language models via probabilistic extraction.Jamie Hayes, Marika Swanberg, Harsh Chaudhari, Itay Yona, Ilia Shumailov, Milad Nasr, Christopher A. Choquette-Choo, Katherine Lee, A. Feder Cooper
2024AAAIArbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification.A. Feder Cooper, Katherine Lee, Madiha Zahrah Choksi, Solon Barocas, Christopher De Sa, James Grimmelmann, Jon M. Kleinberg, Siddhartha Sen, Baobao Zhang
2024ICMLStealing part of a production language model.Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, Florian Tramr
2024NAACLA Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, Daphne Ippolito
2023ICLRQuantifying Memorization Across Neural Language Models.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramr, Chiyuan Zhang
2023ICLRMeasuring Forgetting of Memorized Training Examples.Matthew Jagielski, Om Thakkar, Florian Tramr, Daphne Ippolito, Katherine Lee, Nicholas Carlini, Eric Wallace, Shuang Song, Abhradeep Guha Thakurta, Nicolas Papernot, Chiyuan Zhang
2023INLGReverse-Engineering Decoding Strategies Given Blackbox Access to a Language Generation System.Daphne Ippolito, Nicholas Carlini, Katherine Lee, Milad Nasr, Yun William Yu
2023INLGPreventing Generation of Verbatim Memorization in Language Models Gives a False Sense of Privacy.Daphne Ippolito, Florian Tramr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, Nicholas Carlini
2022ACLDeduplicating Training Data Makes Language Models Better.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini
2021AMIAPredictive Modeling of Healthcare Utilization Metrics Identifies Adult Patients at High Risk for Suicide Attempt in the Primary Care Setting.Katherine Lee, Colin G. Walsh
2004COLINGAnalysis and Detection of Reading Miscues for Interactive Literacy Tutors.Katherine Lee, Andreas Hagen, Nicholas Romanyshyn, Sean Martin, Bryan L. Pellom