| 2026 | ACL | How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects. | Leonardo Bertolazzi, Sandro Pezzelle, Raffaella Bernardi |
| 2026 | ACL | Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance. | Yunchong Huang, Gianni Barlacchi, Sandro Pezzelle |
| 2026 | EACL | Vision-Language Models Align with Human Neural Representations in Concept Processing. | Anna Bavaresco, Marianne de Heer Kloots, Sandro Pezzelle, Raquel Fernndez |
| 2026 | EACL | Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models. | Kumiko Nakajima, Jan Zuiderveld, Sandro Pezzelle |
| 2025 | ACL | LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks. | Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernndez, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andr F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K. Surikuchi, Ece Takmaz, Alberto Testoni |
| 2025 | ACL | They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse. | Walter Paci, Alessandro Panunzi, Sandro Pezzelle |
| 2025 | ACL | From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions. | Nathanal Carraz Rakotonirina, Mohammed Hamdy, Jon Ander Campos, Lucas Weber, Alberto Testoni, Marzieh Fadaee, Sandro Pezzelle, Marco Del Tredici |
| 2025 | COLING | If I feel smart, I will do the right thing: Combining Complementary Multimodal Information in Visual Language Models. | Yuyu Bai, Sandro Pezzelle |
| 2024 | ACL | Naming, Describing, and Quantifying Visual Objects in Humans and LLMs. | Alberto Testoni, Juell Sprott, Sandro Pezzelle |
| 2024 | ACL | Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST! | Frank Wildenburg, Michael Hanna, Sandro Pezzelle |
| 2024 | EACL | Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes. | Ece Takmaz, Sandro Pezzelle, Raquel Fernndez |
| 2024 | EMNLP | Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition. | Aditya Kaushik Surikuchi, Raquel Fernndez, Sandro Pezzelle |
| 2023 | ACL | Dealing with Semantic Underspecification in Multimodal NLP. | Sandro Pezzelle |
| 2023 | ACL | Speaking the Language of Your Listener: Audience-Aware Adaptation via Plug-and-Play Theory of Mind. | Ece Takmaz, Nicolo' Brandizzi, Mario Giulianelli, Sandro Pezzelle, Raquel Fernndez |
| 2023 | EACL | A Psycholinguistic Analysis of BERT's Representations of Compounds. | Lars Buijtelaar, Sandro Pezzelle |
| 2023 | EMNLP | When Language Models Fall in Love: Animacy Processing in Transformer Language Models. | Michael Hanna, Yonatan Belinkov, Sandro Pezzelle |
| 2023 | EMNLP | The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models. | Xinyi Chen, Raquel Fernndez, Sandro Pezzelle |
| 2023 | EMNLP | GROOViST: A Metric for Grounding Objects in Visual Storytelling. | Aditya K. Surikuchi, Sandro Pezzelle, Raquel Fernndez |
| 2022 | CogSci | Time Alignment between Gaze and Speech in Image Descriptions: Exploring Theories of Linearization. | Ece Takmaz, Sandro Pezzelle, Raquel Fernndez |
| 2021 | NAACL | EaSe: A Diagnostic Tool for VQA based on Answer Diversity. | Shailza Jolly, Sandro Pezzelle, Moin Nabi |
| 2020 | CogSci | Asking questions with a big impact: Adapting to other interpretations of gradable adjectives. | Sandro Pezzelle, Raquel Fernndez |
| 2020 | EMNLP | Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision. | Sandro Pezzelle, Claudio Greco, Greta Gandolfi, Eleonora Gualdoni, Raffaella Bernardi |
| 2020 | EMNLP | Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational Contexts. | Ece Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair, Raquel Fernndez |
| 2020 | EMNLP | Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze. | Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernndez |
| 2019 | EMNLP | Is the Red Square Big? MALeViC: Modeling Adjectives Leveraging Visual Contexts. | Sandro Pezzelle, Raquel Fernndez |
| 2019 | EMNLP | Big Generalizations with Small Data: Exploring the Role of Training Samples in Learning Adjectives of Size. | Sandro Pezzelle, Raquel Fernndez |
| 2018 | ACL | Some of Them Can be Guessed! Exploring the Effect of Linguistic Context in Predicting Quantifiers. | Sandro Pezzelle, Shane Steinert-Threlkeld, Raffaella Bernardi, Jakub Szymanik |
| 2018 | NAACL | Comparatives, Quantifiers, Proportions: a Multi-Task Model for the Learning of Quantities from Vision. | Sandro Pezzelle, Ionut-Teodor Sorodoc, Raffaella Bernardi |
| 2017 | ACL | FOIL it! Find One mismatch between Image and Language caption. | Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurlie Herbelot, Moin Nabi, Enver Sangineto, Raffaella Bernardi |
| 2017 | EACL | Be Precise or Fuzzy: Learning the Meaning of Cardinals and Quantifiers from Vision. | Sandro Pezzelle, Marco Marelli, Raffaella Bernardi |
| 2016 | ACL | The LAMBADA dataset: Word prediction requiring a broad discourse context. | Denis Paperno, Germn Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernndez |
| 2016 | ACL | Building a Bagpipe with a Bag and a Pipe: Exploring Conceptual Combination in Vision. | Sandro Pezzelle, Ravi Shekhar, Raffaella Bernardi |
| 2016 | ACL | "Look, some Green Circles!": Learning to Quantify from Images. | Ionut Sorodoc, Angeliki Lazaridou, Gemma Boleda, Aurlie Herbelot, Sandro Pezzelle, Raffaella Bernardi |