Skip to content

Yonatan Belinkov

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

82

Venues

15

Active years

2013–2026

Best venue rank

A*

Where they publish

Papers

82 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLCRISP: Persistent Concept Unlearning via Sparse Autoencoders.Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
2026ACLMasked by Consensus: Disentangling Privileged Knowledge in LLM Correctness.Tomer Ashuach, Shai Gretz, Yoav Katz, Yonatan Belinkov, Liat Ein-Dor
2026ACLFollow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models.Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz
2026ACLWill it Merge? On The Causes of Model Mergeability.Adir Rahamim, Asaf Yehudai, Boaz Carmeli, Leshem Choshen, Yosi Mass, Yonatan Belinkov
2026ACLMechanisms of Prompt-Induced Hallucination in Vision-Language Models.William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald
2025AAAIUnsupervised Translation of Emergent Communication.Ido Levy, Orr Paradise, Boaz Carmeli, Ron Meir, Shafi Goldwasser, Yonatan Belinkov
2025ACLREVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space.Tomer Ashuach, Martin Tutek, Yonatan Belinkov
2025ACLPosition-aware Automatic Circuit Discovery.Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov
2025EMNLPSAEs Are Good for Steering - If You Select the Right Features.Dana Arad, Aaron Mueller, Yonatan Belinkov
2025EMNLPTrust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer.Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov
2025EMNLPMeasuring Chain of Thought Faithfulness by Unlearning Reasoning Steps.Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov
2025EMNLPBack Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models.Zeping Yu, Yonatan Belinkov, Sophia Ananiadou
2025ICLRCtD: Composition through Decomposition in Emergent Communication.Boaz Carmeli, Ron Meir, Yonatan Belinkov
2025ICLRSparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
2025ICLRArithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics.Yaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan Belinkov
2025ICLRLLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations.Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, Yonatan Belinkov
2025ICLRAnswer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions.Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, Ashish Sabharwal
2025ICMLMIB: A Mechanistic Interpretability Benchmark.Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Ivn Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov
2025NAACLPadding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models.Michael Toker, Ido Galil, Hadas Orgad, Rinon Gal, Yoad Tewel, Gal Chechik, Yonatan Belinkov
2025SIGIRReverse-Engineering the Retrieval Process in GenIR Models.Anja Reusch, Yonatan Belinkov
2024AAAIAccelerating the Global Aggregation of Local Explanations.Alon Mor, Yonatan Belinkov, Benny Kimelfeld
2024ACLConcept-Best-Matching: Evaluating Compositionality In Emergent Communication.Boaz Carmeli, Yonatan Belinkov, Ron Meir
2024ACLDiffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines.Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov
2024EACLGenerating Benchmarks for Factuality Evaluation of Language Models.Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, Yoav Shoham
2024EACLA Dataset for Metaphor Detection in Early Medieval Hebrew Poetry.Michael Toker, Oren Mishali, Ophir Mnz-Manor, Benny Kimelfeld, Yonatan Belinkov
2024EMNLPBackward Lens: Projecting Language Model Gradients into the Vocabulary Space.Shahar Katz, Yonatan Belinkov, Mor Geva, Lior Wolf
2024EMNLPFast Forwarding Low-Rank Training.Adir Rahamim, Naomi Saphra, Sara Kangaslahti, Yonatan Belinkov
2024ICLRLinearity of Relation Decoding in Transformer Language Models.Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, David Bau
2024ICLRFine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking.Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, David Bau
2024NAACLReFACT: Updating Text-to-Image Models by Editing the Text Encoder.Dana Arad, Hadas Orgad, Yonatan Belinkov
2024NAACLLeveraging Prototypical Representations for Mitigating Social Bias without Demographic Information.Shadi Iskander, Kira Radinsky, Yonatan Belinkov
2024NAACLContraSim - Analyzing Neural Representations Based on Contrastive Learning.Adir Rahamim, Yonatan Belinkov
2024WACVUnified Concept Editing in Diffusion Models.Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzynska, David Bau
2023AAAIEmergent Quantized Communication.Boaz Carmeli, Ron Meir, Yonatan Belinkov
2023ACLShielded Representations: Protecting Sensitive Attributes Through Iterative Gradient-Based Projection.Shadi Iskander, Kira Radinsky, Yonatan Belinkov
2023ACLBLIND: Bias Removal With No Demographics.Hadas Orgad, Yonatan Belinkov
2023ACLWhat Are You Token About? Dense Retrieval as Distributions Over the Vocabulary.Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson
2023ACLParallel Context Windows for Large Language Models.Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham
2023EMNLPWhen Language Models Fall in Love: Animacy Processing in Transformer Language Models.Michael Hanna, Yonatan Belinkov, Sandro Pezzelle
2023EMNLPVISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers.Shahar Katz, Yonatan Belinkov
2023EMNLPA Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.Alessandro Stolfo, Yonatan Belinkov, Mrinmaya Sachan
2023ICCVEditing Implicit Assumptions in Text-to-Image Diffusion Models.Hadas Orgad, Bahjat Kawar, Yonatan Belinkov
2023ICLRMultiple sequence alignment as a sequence-to-sequence learning problem.Edo Dotan, Yonatan Belinkov, Oren Avram, Elya Wygoda, Noa Ecker, Michael Alburquerque, Omri Keren, Gil Loewenthal, Tal Pupko
2023ICLRMass-Editing Memory in a Transformer.Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, David Bau
2022AAAISupervising Model Attention with Human Explanations for Robust Natural Language Inference.Joe Stacey, Yonatan Belinkov, Marek Rei
2022EMNLPA Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference.Kerem Zaman, Yonatan Belinkov
2022ICLROn the Pitfalls of Analyzing Individual Neurons in Language Models.Omer Antverg, Yonatan Belinkov
2022NAACLHow Gender Debiasing Affects Internal Model Representations, and Why It Matters.Hadas Orgad, Seraphina Goldfarb-Tarrant, Yonatan Belinkov
2021ACLCausal Analysis of Syntactic Agreement Mechanisms in Neural Language Models.Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov
2021EACLProbing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?Abhilasha Ravichander, Yonatan Belinkov, Eduard H. Hovy
2021EMNLPDebiasing Methods in Natural Language Understanding Make Bias More Accessible.Michael Mendelson, Yonatan Belinkov
2021ICASSPSimilarity Analysis of Self-Supervised Speech Representations.Yu-An Chung, Yonatan Belinkov, James R. Glass
2021ICLRVariational Information Bottleneck for Effective Low-Resource Fine-Tuning.Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson
2021ICLRLearning from others' mistakes: Avoiding dataset biases without modeling them.Victor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. Rush
2020ACLThe Sensitivity of Language Models and Humans to Winograd Schema Perturbations.Mostafa Abdou, Vinit Ravishankar, Maria Barrett, Yonatan Belinkov, Desmond Elliott, Anders Sgaard
2020ACLInterpretability and Analysis in Neural NLP.Yonatan Belinkov, Sebastian Gehrmann, Ellie Pavlick
2020ACLEnd-to-End Bias Mitigation by Modelling Biases in Corpora.Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson
2020ACLSimilarity Analysis of Contextual Word Representation Models.John M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass
2020EMNLPAnalyzing Redundancy in Pretrained Transformer Models.Fahim Dalvi, Hassan Sajjad, Nadir Durrani, Yonatan Belinkov
2020EMNLPAnalyzing Individual Neurons in Pre-trained Language Models.Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan Belinkov
2020ICLRA Constructive Prediction of the Generalization Error Across Scales.Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir Shavit
2019AAAIWhat Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models.Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, James R. Glass
2019AAAINeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks.Fahim Dalvi, Avery Nortonsmith, Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, James R. Glass
2019ACLDon't Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference.Yonatan Belinkov, Adam Poliak, Stuart M. Shieber, Benjamin Van Durme, Alexander M. Rush
2019ACLImproving Neural Language Models by Segmenting, Attending, and Predicting the Future.Hongyin Luo, Lan Jiang, Yonatan Belinkov, James R. Glass
2019CogSciCharacter-based Surprisal as a Model of Reading Difficulty in the Presence of Errors.Michael Hahn, Frank Keller, Yonatan Bisk, Yonatan Belinkov
2019ICLRIdentifying and Controlling Important Neurons in Neural Machine Translation.Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass
2019InterspeechAnalyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech Recognition.Yonatan Belinkov, Ahmed Ali, James R. Glass
2019NAACLAnalysis Methods in Neural Language Processing: A Survey.Yonatan Belinkov, James R. Glass
2019NAACLOne Size Does Not Fit All: Comparing NMT Representations of Different Granularities.Nadir Durrani, Fahim Dalvi, Hassan Sajjad, Yonatan Belinkov, Preslav Nakov
2019NAACLLinguistic Knowledge and Transferability of Contextual Representations.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, Noah A. Smith
2018ICLRSynthetic and Natural Noise Both Break Neural Machine Translation.Yonatan Belinkov, Yonatan Bisk
2018NAACLOn the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference.Adam Poliak, Yonatan Belinkov, James R. Glass, Benjamin Van Durme
2017ACLWhat do Neural Machine Translation Models Learn about Morphology?Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, James R. Glass
2017ACLChallenging Language-Dependent Segmentation for Arabic: An Application to Machine Translation and Part-of-Speech Tagging.Hassan Sajjad, Fahim Dalvi, Nadir Durrani, Ahmed Abdelali, Yonatan Belinkov, Stephan Vogel
2017ICLRFine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks.Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, Yoav Goldberg
2017IJCNLPEvaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks.Yonatan Belinkov, Llus Mrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass
2017IJCNLPUnderstanding and Improving Morphological Learning in the Neural Machine Translation Decoder.Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Stephan Vogel
2017InterspeechQMDIS: QCRI-MIT Advanced Dialect Identification System.Sameer Khurana, Maryam Najafian, Ahmed Ali, Tuka Al Hanai, Yonatan Belinkov, James R. Glass
2016COLINGNeural Attention for Learning to Rank Questions in Community Question Answering.Salvatore Romeo, Giovanni Da San Martino, Alberto Barrn-Cedeo, Alessandro Moschitti, Yonatan Belinkov, Wei-Ning Hsu, Yu Zhang, Mitra Mohtarami, James R. Glass
2015EMNLPArabic Diacritization with Recurrent Neural Networks.Yonatan Belinkov, James R. Glass
2013ACLTranslating Dialectal Arabic to English.Hassan Sajjad, Kareem Darwish, Yonatan Belinkov