| 2026 | LREC | Towards a Diagnostic and Predictive Evaluation Methodology for Sequence Labeling Tasks. | Elena lvarez Mellado, Julio Gonzalo |
| 2025 | ACL | Evaluating Sequence Labeling on the basis of Information Theory. | Enrique Amig, Elena lvarez Mellado, Julio Gonzalo, Jorge Carrillo-de-Albornoz |
| 2025 | ACL | The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing. | Guillermo Marco, Julio Gonzalo, Vctor Fresno |
| 2025 | COLING | Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs. | Guillermo Marco, Luz Rello, Julio Gonzalo |
| 2025 | COLING | Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination. | Eva Snchez-Salido, Roser Morante, Julio Gonzalo, Guillermo Marco, Jorge Carrillo-de-Albornoz, Laura Plaza, Enrique Amig, Andrs Fernndez Garca, Alejandro Benito-Santos, Adrin Ghajari Espinosa, Vctor Fresno |
| 2025 | ECIR | EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos. | Laura Plaza, Jorge Carrillo-de-Albornoz, Ivn rcos, Paolo Rosso, Damiano Spina, Enrique Amig, Julio Gonzalo, Roser Morante |
| 2024 | COLING | A Web Portal about the State of the Art of NLP Tasks in Spanish. | Enrique Amig, Jorge Carrillo-de-Albornoz, Andrs Fernndez, Julio Gonzalo, Guillermo Marco, Roser Morante, Laura Plaza, Jacobo Pedrosa |
| 2024 | ECIR | The CLEF 2024 Monster Track: One Lab to Rule Them All. | Nicola Ferro, Julio Gonzalo, Jussi Karlgren, Henning Mller |
| 2024 | ECIR | EXIST 2024: sEXism Identification in Social neTworks and Memes. | Laura Plaza, Jorge Carrillo-de-Albornoz, Enrique Amig, Julio Gonzalo, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Vctor Ruiz |
| 2024 | EMNLP | Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing? | Guillermo Marco, Julio Gonzalo, Mara Teresa Mateo Girona, Ramn Santos |
| 2023 | ECIR | Overview of EXIST 2023: sEXism Identification in Social NeTworks. | Laura Plaza, Jorge Carrillo-de-Albornoz, Roser Morante, Enrique Amig, Julio Gonzalo, Damiano Spina, Paolo Rosso |
| 2020 | ACL | An Effectiveness Metric for Ordinal Classification: Formal Properties and Experimental Results. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro, Jorge Carrillo-de-Albornoz |
| 2017 | ECIR | A Formal and Empirical Study of Unsupervised Signal Combination for Textual Similarity Tasks. | Enrique Amig, Fernando Giner, Julio Gonzalo, Felisa Verdejo |
| 2017 | ECIR | Sentiment Propagation for Predicting Reputation Polarity. | Anastasia Giachanou, Julio Gonzalo, Ida Mele, Fabio Crestani |
| 2017 | SIGIR | EvALL: Open Access Evaluation for Information Access Systems. | Enrique Amig, Jorge Carrillo de Albornoz, Mario Almagro-Cdiz, Julio Gonzalo, Javier Rodrguez-Vidal, Felisa Verdejo |
| 2016 | ECIR | Tweet Stream Summarization for Online Reputation Management. | Jorge Carrillo de Albornoz, Enrique Amig, Laura Plaza, Julio Gonzalo |
| 2016 | ECIR | Monitoring Reputation in the Wild Online West. | Julio Gonzalo |
| 2015 | ECIR | A Formal Approach to Effectiveness Metrics for Information Access: Retrieval, Filtering, and Clustering. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro |
| 2014 | ECIR | ORMA: A Semi-automatic Tool for Online Reputation Monitoring in Twitter. | Jorge Carrillo de Albornoz, Enrique Amig, Damiano Spina, Julio Gonzalo |
| 2014 | SIGIR | A general account of effectiveness metrics for information tasks: retrieval, filtering, and clustering. | Enrique Amig, Julio Gonzalo, Stefano Mizzaro |
| 2014 | SIGIR | SIGIR 2014 workshop on semantic matching in information retrieval. | Julio Gonzalo, Hang Li, Alessandro Moschitti, Jun Xu |
| 2014 | SIGIR | Learning similarity functions for topic detection in online reputation monitoring. | Damiano Spina, Julio Gonzalo, Enrique Amig |
| 2013 | CIKM | An unsupervised transfer learning approach to discover topics for online reputation management. | Tamara Martn-Wanton, Julio Gonzalo, Enrique Amig |
| 2013 | SIGIR | A general evaluation measure for document organization tasks. | Enrique Amig, Julio Gonzalo, Felisa Verdejo |
| 2012 | COLING | Automatic Extraction of Polar Adjectives for the Creation of Polarity Lexicons. | Silvia Vzquez, Muntsa Padr, Nria Bel, Julio Gonzalo |
| 2012 | NAACL | The Heterogeneity Principle in Evaluation Measures for Automatic Summarization. | Enrique Amig, Julio Gonzalo, Felisa Verdejo |
| 2011 | EMNLP | Corroborating Text Evaluation Results with Heterogeneous Measures. | Enrique Amig, Julio Gonzalo, Jess Gimnez, Felisa Verdejo |
| 2010 | ACL | Wikipedia as Sense Inventory to Improve Diversity in Web Search Results. | Celina Santamara, Julio Gonzalo, Javier Artiles |
| 2009 | ACL | The Contribution of Linguistic Features to Automatic Machine Translation Evaluation. | Enrique Amig, Jess Gimnez, Julio Gonzalo, Felisa Verdejo |
| 2009 | ACL | The Impact of Query Refinement in the Web People Search Task. | Javier Artiles, Julio Gonzalo, Enrique Amig |
| 2009 | EMNLP | The role of named entities in Web People Search. | Javier Artiles, Enrique Amig, Julio Gonzalo |
| 2008 | ECIR | Workshop on Novel Methodologies for Evaluation in Information Retrieval. | Mark Sanderson, Martin Braschler, Nicola Ferro, Julio Gonzalo |
| 2008 | LREC | From Research to Application in Multilingual Information Access: the Contribution of Evaluation. | Carol Peters, Martin Braschler, Giorgio Maria Di Nunzio, Nicola Ferro, Julio Gonzalo, Mark Sanderson |
| 2008 | WWW | Web people search: results of the first evaluation and the plan for the second. | Javier Artiles, Satoshi Sekine, Julio Gonzalo |
| 2006 | ACL | MT Evaluation: Human-Like vs. Human Acceptable. | Enrique Amig, Jess Gimnez, Julio Gonzalo, Llus Mrquez |
| 2005 | ACL | QARLA: A Framework for the Evaluation of Text Summarization Systems. | Enrique Amig, Julio Gonzalo, Anselmo Peas, Felisa Verdejo |
| 2005 | ACL | Evaluating DUC 2004 Tasks with the QARLA Framework. | Enrique Amig, Julio Gonzalo, Anselmo Peas, Felisa Verdejo |
| 2005 | ICFCA | Automatic Selection of Noun Phrases as Document Descriptors in an FCA-Based Information Retrieval System. | Juan M. Cigarrn, Anselmo Peas, Julio Gonzalo, Felisa Verdejo |
| 2005 | SIGIR | A testbed for people searching strategies in the WWW. | Javier Artiles, Julio Gonzalo, Felisa Verdejo |
| 2005 | SIGIR | The impact of evaluation on multilingual text retrieval. | Julio Gonzalo, Carol Peters |
| 2005 | SPIRE | Evaluating Hierarchical Clustering of Search Results. | Juan M. Cigarrn, Anselmo Peas, Julio Gonzalo, Felisa Verdejo |
| 2004 | ACL | An Empirical Study of Information Synthesis Task. | Enrique Amig, Julio Gonzalo, Vctor Peinado, Anselmo Peas, Felisa Verdejo |
| 2004 | COLING | Using syntactic information to extract relevant terms for multi-document summarization. | Enrique Amig, Julio Gonzalo, Vctor Peinado, Anselmo Peas, Felisa Verdejo |
| 2004 | ICFCA | Browsing Search Results via Formal Concept Analysis: Automatic Selection of Attributes. | Juan M. Cigarrn, Julio Gonzalo, Anselmo Peas, Felisa Verdejo |
| 2004 | LREC | The Future of Evaluation for Cross-Language Information Retrieval Systems. | Carol Peters, Martin Braschler, Khalid Choukri, Julio Gonzalo, Michael Kluck |
| 2003 | CICLING | Suggesting Named Entities for Information Access. | Enrique Amig, Anselmo Peas, Julio Gonzalo, Felisa Verdejo |
| 2001 | NLDB | The Role of Conceptual Relation in Word Sense Disambiguation. | David Fernndez-Amors, Julio Gonzalo, Felisa Verdejo |
| 2001 | NLDB | Cross-Language Information Access through Phrase Browsing. | Anselmo Peas, Julio Gonzalo, Felisa Verdejo |
| 2000 | LREC | Evaluating Wordnets in Cross-language Information Retrieval: the ITEM Search Engine. | Felisa Verdejo, Julio Gonzalo, Anselmo Peas, Fernando Lpez-Ostenero, David Fernndez-Amors |
| 1999 | EACL | An Open Distance Learning Web-Course for NLP in IR. | Felisa Verdejo, Julio Gonzalo, Anselmo Peas |
| 1999 | EMNLP | Lexical ambiguity and Information Retrieval revisited. | Julio Gonzalo, Anselmo Peas, Felisa Verdejo |
| 1998 | COLING | Indexing with WordNet synsets can improve text retrieval. | Julio Gonzalo, Felisa Verdejo, Irina Chugur, Juan M. Cigarrn |