Skip to content

Fazl Barez

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

15

Venues

4

Active years

2023–2026

Best venue rank

A*

Where they publish

Papers

15 indexed papers, newest first.

YearVenueTitleAuthors
2026ACLMake Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing.Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi, Philip Quirke, Austin Meek, Fazl Barez, Amir Abdullah
2025EMNLPSame Question, Different Words: A Latent Adversarial Framework for Prompt Robustness.Tingchen Fu, Fazl Barez
2025EMNLPPrecise In-Parameter Concept Erasure in Large Language Models.Yoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez, Mor Geva
2025EMNLPBeyond Linear Steering: Unified Multi-Attribute Control for Language Models.Narmeen Oozeer, Luke Marks, Fazl Barez, Amir Abdullah
2025EMNLPTrust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer.Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov
2025ICLRTowards Interpreting Visual Information Processing in Vision-Language Models.Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez
2025ICMLPoisonBench: Assessing Language Model Vulnerability to Poisoned Preference Data.Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen, David Krueger, Fazl Barez
2024ACLLarge Language Models Relearn Removed Concepts.Michelle Lo, Fazl Barez, Shay B. Cohen
2024EMNLPTowards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models.Michael Lan, Philip Torr, Fazl Barez
2024EMNLPInterpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions.Clement Neo, Shay B. Cohen, Fazl Barez
2024ICLRUnderstanding Addition in Transformers.Philip Quirke, Fazl Barez
2024ICMLPosition: Near to Mid-term Risks and Opportunities of Open-Source Generative AI.Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schrder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Rttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster
2024ICMLValue-Evolutionary-Based Reinforcement Learning.Pengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng, Fazl Barez
2023ACLThe Larger they are, the Harder they Fail: Language Models do not Recognize Identifier Swaps in Python.Antonio Valerio Miceli Barone, Fazl Barez, Shay B. Cohen, Ioannis Konstas
2023ACLDetecting Edit Failures In Large Language Models: An Improved Specificity Benchmark.Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, Fazl Barez