| 2026 | ACL | CRISP: Persistent Concept Unlearning via Sparse Autoencoders. | Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov |
| 2026 | ACL | Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness. | Tomer Ashuach, Shai Gretz, Yoav Katz, Yonatan Belinkov, Liat Ein-Dor |
| 2026 | ACL | Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models. | Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz |
| 2026 | ACL | Will it Merge? On The Causes of Model Mergeability. | Adir Rahamim, Asaf Yehudai, Boaz Carmeli, Leshem Choshen, Yosi Mass, Yonatan Belinkov |
| 2026 | ACL | Mechanisms of Prompt-Induced Hallucination in Vision-Language Models. | William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald |
| 2025 | AAAI | Unsupervised Translation of Emergent Communication. | Ido Levy, Orr Paradise, Boaz Carmeli, Ron Meir, Shafi Goldwasser, Yonatan Belinkov |
| 2025 | ACL | REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space. | Tomer Ashuach, Martin Tutek, Yonatan Belinkov |
| 2025 | ACL | Position-aware Automatic Circuit Discovery. | Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov |
| 2025 | EMNLP | SAEs Are Good for Steering - If You Select the Right Features. | Dana Arad, Aaron Mueller, Yonatan Belinkov |
| 2025 | EMNLP | Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer. | Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov |
| 2025 | EMNLP | Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps. | Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov |
| 2025 | EMNLP | Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models. | Zeping Yu, Yonatan Belinkov, Sophia Ananiadou |
| 2025 | ICLR | CtD: Composition through Decomposition in Emergent Communication. | Boaz Carmeli, Ron Meir, Yonatan Belinkov |
| 2025 | ICLR | Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. | Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller |
| 2025 | ICLR | Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics. | Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov |
| 2025 | ICLR | LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. | Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, Yonatan Belinkov |
| 2025 | ICLR | Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions. | Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, Ashish Sabharwal |
| 2025 | ICML | MIB: A Mechanistic Interpretability Benchmark. | Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Ivn Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov |
| 2025 | NAACL | Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models. | Michael Toker, Ido Galil, Hadas Orgad, Rinon Gal, Yoad Tewel, Gal Chechik, Yonatan Belinkov |
| 2025 | SIGIR | Reverse-Engineering the Retrieval Process in GenIR Models. | Anja Reusch, Yonatan Belinkov |
| 2024 | AAAI | Accelerating the Global Aggregation of Local Explanations. | Alon Mor, Yonatan Belinkov, Benny Kimelfeld |
| 2024 | ACL | Concept-Best-Matching: Evaluating Compositionality In Emergent Communication. | Boaz Carmeli, Yonatan Belinkov, Ron Meir |
| 2024 | ACL | Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines. | Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov |
| 2024 | EACL | Generating Benchmarks for Factuality Evaluation of Language Models. | Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, Yoav Shoham |
| 2024 | EACL | A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry. | Michael Toker, Oren Mishali, Ophir Mnz-Manor, Benny Kimelfeld, Yonatan Belinkov |
| 2024 | EMNLP | Backward Lens: Projecting Language Model Gradients into the Vocabulary Space. | Shahar Katz, Yonatan Belinkov, Mor Geva, Lior Wolf |
| 2024 | EMNLP | Fast Forwarding Low-Rank Training. | Adir Rahamim, Naomi Saphra, Sara Kangaslahti, Yonatan Belinkov |
| 2024 | ICLR | Linearity of Relation Decoding in Transformer Language Models. | Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, David Bau |
| 2024 | ICLR | Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking. | Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, David Bau |
| 2024 | NAACL | ReFACT: Updating Text-to-Image Models by Editing the Text Encoder. | Dana Arad, Hadas Orgad, Yonatan Belinkov |
| 2024 | NAACL | Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information. | Shadi Iskander, Kira Radinsky, Yonatan Belinkov |
| 2024 | NAACL | ContraSim - Analyzing Neural Representations Based on Contrastive Learning. | Adir Rahamim, Yonatan Belinkov |
| 2024 | WACV | Unified Concept Editing in Diffusion Models. | Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzynska, David Bau |
| 2023 | AAAI | Emergent Quantized Communication. | Boaz Carmeli, Ron Meir, Yonatan Belinkov |
| 2023 | ACL | Shielded Representations: Protecting Sensitive Attributes Through Iterative Gradient-Based Projection. | Shadi Iskander, Kira Radinsky, Yonatan Belinkov |
| 2023 | ACL | BLIND: Bias Removal With No Demographics. | Hadas Orgad, Yonatan Belinkov |
| 2023 | ACL | What Are You Token About? Dense Retrieval as Distributions Over the Vocabulary. | Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson |
| 2023 | ACL | Parallel Context Windows for Large Language Models. | Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham |
| 2023 | EMNLP | When Language Models Fall in Love: Animacy Processing in Transformer Language Models. | Michael Hanna, Yonatan Belinkov, Sandro Pezzelle |
| 2023 | EMNLP | VISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers. | Shahar Katz, Yonatan Belinkov |
| 2023 | EMNLP | A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis. | Alessandro Stolfo, Yonatan Belinkov, Mrinmaya Sachan |
| 2023 | ICCV | Editing Implicit Assumptions in Text-to-Image Diffusion Models. | Hadas Orgad, Bahjat Kawar, Yonatan Belinkov |
| 2023 | ICLR | Multiple sequence alignment as a sequence-to-sequence learning problem. | Edo Dotan, Yonatan Belinkov, Oren Avram, Elya Wygoda, Noa Ecker, Michael Alburquerque, Omri Keren, Gil Loewenthal, Tal Pupko |
| 2023 | ICLR | Mass-Editing Memory in a Transformer. | Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, David Bau |
| 2022 | AAAI | Supervising Model Attention with Human Explanations for Robust Natural Language Inference. | Joe Stacey, Yonatan Belinkov, Marek Rei |
| 2022 | EMNLP | A Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference. | Kerem Zaman, Yonatan Belinkov |
| 2022 | ICLR | On the Pitfalls of Analyzing Individual Neurons in Language Models. | Omer Antverg, Yonatan Belinkov |
| 2022 | NAACL | How Gender Debiasing Affects Internal Model Representations, and Why It Matters. | Hadas Orgad, Seraphina Goldfarb-Tarrant, Yonatan Belinkov |
| 2021 | ACL | Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models. | Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov |
| 2021 | EACL | Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance? | Abhilasha Ravichander, Yonatan Belinkov, Eduard H. Hovy |
| 2021 | EMNLP | Debiasing Methods in Natural Language Understanding Make Bias More Accessible. | Michael Mendelson, Yonatan Belinkov |
| 2021 | ICASSP | Similarity Analysis of Self-Supervised Speech Representations. | Yu-An Chung, Yonatan Belinkov, James R. Glass |
| 2021 | ICLR | Variational Information Bottleneck for Effective Low-Resource Fine-Tuning. | Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson |
| 2021 | ICLR | Learning from others' mistakes: Avoiding dataset biases without modeling them. | Victor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. Rush |
| 2020 | ACL | The Sensitivity of Language Models and Humans to Winograd Schema Perturbations. | Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, Anders Sgaard |
| 2020 | ACL | Interpretability and Analysis in Neural NLP. | Yonatan Belinkov, Sebastian Gehrmann, Ellie Pavlick |
| 2020 | ACL | End-to-End Bias Mitigation by Modelling Biases in Corpora. | Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson |
| 2020 | ACL | Similarity Analysis of Contextual Word Representation Models. | John M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2020 | EMNLP | Analyzing Redundancy in Pretrained Transformer Models. | Fahim Dalvi, Hassan Sajjad, Nadir Durrani, Yonatan Belinkov |
| 2020 | EMNLP | Analyzing Individual Neurons in Pre-trained Language Models. | Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan Belinkov |
| 2020 | ICLR | A Constructive Prediction of the Generalization Error Across Scales. | Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir Shavit |
| 2019 | AAAI | What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models. | Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, James R. Glass |
| 2019 | AAAI | NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks. | Fahim Dalvi, Avery Nortonsmith, Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, James R. Glass |
| 2019 | ACL | Don't Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference. | Yonatan Belinkov, Adam Poliak, Stuart M. Shieber, Benjamin Van Durme, Alexander M. Rush |
| 2019 | ACL | Improving Neural Language Models by Segmenting, Attending, and Predicting the Future. | Hongyin Luo, Lan Jiang, Yonatan Belinkov, James R. Glass |
| 2019 | CogSci | Character-based Surprisal as a Model of Reading Difficulty in the Presence of Errors. | Michael Hahn, Frank Keller, Yonatan Bisk, Yonatan Belinkov |
| 2019 | ICLR | Identifying and Controlling Important Neurons in Neural Machine Translation. | Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2019 | Interspeech | Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech Recognition. | Yonatan Belinkov, Ahmed Ali, James R. Glass |
| 2019 | NAACL | Analysis Methods in Neural Language Processing: A Survey. | Yonatan Belinkov, James R. Glass |
| 2019 | NAACL | One Size Does Not Fit All: Comparing NMT Representations of Different Granularities. | Nadir Durrani, Fahim Dalvi, Hassan Sajjad, Yonatan Belinkov, Preslav Nakov |
| 2019 | NAACL | Linguistic Knowledge and Transferability of Contextual Representations. | Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, Noah A. Smith |
| 2018 | ICLR | Synthetic and Natural Noise Both Break Neural Machine Translation. | Yonatan Belinkov, Yonatan Bisk |
| 2018 | NAACL | On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference. | Adam Poliak, Yonatan Belinkov, James R. Glass, Benjamin Van Durme |
| 2017 | ACL | What do Neural Machine Translation Models Learn about Morphology? | Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, James R. Glass |
| 2017 | ACL | Challenging Language-Dependent Segmentation for Arabic: An Application to Machine Translation and Part-of-Speech Tagging. | Hassan Sajjad, Fahim Dalvi, Nadir Durrani, Ahmed Abdelali, Yonatan Belinkov, Stephan Vogel |
| 2017 | ICLR | Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks. | Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, Yoav Goldberg |
| 2017 | IJCNLP | Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks. | Yonatan Belinkov, Llus Mrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2017 | IJCNLP | Understanding and Improving Morphological Learning in the Neural Machine Translation Decoder. | Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Stephan Vogel |
| 2017 | Interspeech | QMDIS: QCRI-MIT Advanced Dialect Identification System. | Sameer Khurana, Maryam Najafian, Ahmed Ali, Tuka Al Hanai, Yonatan Belinkov, James R. Glass |
| 2016 | COLING | Neural Attention for Learning to Rank Questions in Community Question Answering. | Salvatore Romeo, Giovanni Da San Martino, Alberto Barrn-Cedeo, Alessandro Moschitti, Yonatan Belinkov, Wei-Ning Hsu, Yu Zhang, Mitra Mohtarami, James R. Glass |
| 2015 | EMNLP | Arabic Diacritization with Recurrent Neural Networks. | Yonatan Belinkov, James R. Glass |
| 2013 | ACL | Translating Dialectal Arabic to English. | Hassan Sajjad, Kareem Darwish, Yonatan Belinkov |