| 2024 | EACL | Anchor Points: Benchmarking Models with Much Fewer Examples. | Rajan Vivek, Kawin Ethayarajh, Diyi Yang, Douwe Kiela |
| 2024 | ICML | Model Alignment as Prospect Theoretic Optimization. | Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, Douwe Kiela |
| 2022 | ACL | Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words. | Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, Dan Jurafsky |
| 2022 | ACL | Richer Countries and Richer Representations. | Kaitlyn Zhou, Kawin Ethayarajh, Dan Jurafsky |
| 2022 | EMNLP | The Authenticity Gap in Human Evaluation. | Kawin Ethayarajh, Dan Jurafsky |
| 2022 | ICML | Understanding Dataset Difficulty with | Kawin Ethayarajh, Yejin Choi, Swabha Swayamdipta |
| 2021 | ACL | Attention Flows are Shapley Value Explanations. | Kawin Ethayarajh, Dan Jurafsky |
| 2021 | EMNLP | Conditional probing: measuring usable information beyond a baseline. | John Hewitt, Kawin Ethayarajh, Percy Liang, Christopher D. Manning |
| 2020 | ACL | Is Your Classifier Actually Biased? Measuring Fairness under Uncertainty with Bernstein Bounds. | Kawin Ethayarajh |
| 2020 | EMNLP | Utility is in the Eye of the User: A Critique of NLP Leaderboards. | Kawin Ethayarajh, Dan Jurafsky |
| 2019 | ACL | Understanding Undesirable Word Embedding Associations. | Kawin Ethayarajh, David Duvenaud, Graeme Hirst |
| 2019 | ACL | Towards Understanding Linear Word Analogies. | Kawin Ethayarajh, David Duvenaud, Graeme Hirst |
| 2019 | EMNLP | How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings. | Kawin Ethayarajh |
| 2019 | EMNLP | Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space. | Kawin Ethayarajh |