| 2026 | AAAI | Control Illusion: The Failure of Instruction Hierarchies in Large Language Models. | Yilin Geng, Haonan Li, Honglin Mu, Xudong Han, Timothy Baldwin, Omri Abend, Eduard H. Hovy, Lea Frermann |
| 2026 | ACL | Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation. | Ekaterina Fadeeva, Aleksandr Rubashevskii, Dzianis Piatrashyn, Roman Vashurin, Shehzaad Dhuliawala, Artem Shelmanov, Timothy Baldwin, Preslav Nakov, Mrinmaya Sachan, Maxim Panov |
| 2026 | ACL | A Multilingual Social Bias Benchmark Incorporating Thinking Processes. | Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin |
| 2026 | ACL | Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models. | Jingwei Ni, Ekaterina Fadeeva, Tianyi Wu, Mubashara Akhtar, Jiaheng Zhang, Elliott Ash, Markus Leippold, Timothy Baldwin, See-Kiong Ng, Artem Shelmanov, Mrinmaya Sachan |
| 2026 | ACL | Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling. | Ivan Rodkin, Daniil Orel, Konstantin Smirnov, Arman Bolatov, Bilal Elbouardi, Besher Hassan, Yuri Kuratov, Aydar Bulatov, Preslav Nakov, Timothy Baldwin, Artem Shelmanov, Mikhail Burtsev |
| 2026 | ACL | ThinkBooster: A Unified Framework for Seamless Test-Time Scaling of LLM Reasoning. | Vladislav Smirnov, Quang-Chieu Nguyen, Sergey Senichev, Minh Ngoc Ta, Ekaterina Fadeeva, Artem Vazhentsev, Daria Galimzianova, Nikolai Rozanov, Viktor Mazanov, Jingwei Ni, Tianyi Wu, Igor Kiselev, Mrinmaya Sachan, Iryna Gurevych, Preslav Nakov, Timothy Baldwin, Artem Shelmanov |
| 2026 | EACL | Do Diacritics Matter? Evaluating the Impact of Arabic Diacritics on Tokenization and LLM Benchmarks. | Go Inoue, Bashar Alhafni, Nizar Habash, Timothy Baldwin |
| 2026 | EACL | On the Interplay between Human Label Variation and Model Fairness. | Kemal Kurniawan, Meladel Mistica, Timothy Baldwin, Jey Han Lau |
| 2026 | EACL | SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning. | Renxi Wang, Honglin Mu, Liqun Ma, Lizhi Lin, Yunlong Feng, Timothy Baldwin, Xudong Han, Haonan Li |
| 2026 | EACL | COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations. | Rui Xing, Preslav Nakov, Timothy Baldwin, Jey Han Lau |
| 2026 | ECIR | Uncertainty Quantification for Large Language Models. | Maxim Panov, Artem Shelmanov, Roman Vashurin, Artem Vazhentsev, Ekaterina Fadeeva, Lyudmila Rvanova, Timothy Baldwin |
| 2025 | ACL | Qorǵau: Evaluating Safety in Kazakh-Russian Bilingual Contexts. | Maiya Goloburda, Nurkhan Laiyk, Diana Turmakhan, Yuxia Wang, Mukhammed Togmanov, Jonibek Mansurov, Askhat Sametov, Nurdaulet Mukhituly, Minghan Wang, Daniil Orel, Zain Muhammad Mujahid, Fajri Koto, Timothy Baldwin, Preslav Nakov |
| 2025 | ACL | Uncertainty Quantification for Large Language Models. | Artem Shelmanov, Maxim Panov, Roman Vashurin, Artem Vazhentsev, Ekaterina Fadeeva, Timothy Baldwin |
| 2025 | COLING | Loki: An Open-Source Tool for Fact Verification. | Haonan Li, Xudong Han, Hao Wang, Yuxia Wang, Minghan Wang, Rui Xing, Yilin Geng, Zenan Zhai, Preslav Nakov, Timothy Baldwin |
| 2025 | COLING | The Gaps between Fine Tuning and In-context Learning in Bias Evaluation and Debiasing. | Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin |
| 2025 | COLING | Does Vision Accelerate Hierarchical Generalization in Neural Language Learners? | Tatsuki Kuribayashi, Timothy Baldwin |
| 2025 | COLING | Human Interest Framing across Cultures: A Case Study on Climate Change. | Gisela Vallejo, Christine de Kock, Timothy Baldwin, Lea Frermann |
| 2025 | EMNLP | Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World. | Saeed Almheiri, Rania Elbadry, Mena Attia, Chenxi Wang, Preslav Nakov, Timothy Baldwin, Fajri Koto |
| 2025 | EMNLP | Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models. | Masahiro Kaneko, Alham Fikri Aji, Timothy Baldwin |
| 2025 | EMNLP | Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection Capabilities. | Masahiro Kaneko, Timothy Baldwin |
| 2025 | EMNLP | BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities. | Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Sara Pieri, Saeed Yahya Alseiari, Shanavas Cholakkal, Khaled Aldahmani, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman H. Khan, Timothy Baldwin, Hisham Cholakkal |
| 2025 | EMNLP | A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs. | Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin |
| 2025 | EMNLP | Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models. | Artem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Gleb Kuzmin, Ivan Lazichny, Alexander Panchenko, Preslav Nakov, Timothy Baldwin, Maxim Panov, Artem Shelmanov |
| 2025 | ICLR | ToolGen: Unified Tool Retrieval and Calling via Generation. | Renxi Wang, Xudong Han, Lei Ji, Shu Wang, Timothy Baldwin, Haonan Li |
| 2025 | IJCAI | An Ethical Dataset from Real-World Interactions Between Users and Large Language Models. | Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin |
| 2025 | IJCNLP | Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization. | Masahiro Kaneko, Zeerak Talat, Timothy Baldwin |
| 2025 | NAACL | Arabic Dataset for LLM Safeguard Evaluation. | Yasser Ashraf, Yuxia Wang, Bin Gu, Preslav Nakov, Timothy Baldwin |
| 2025 | NAACL | Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models. | Gleb Kuzmin, Neemesh Yadav, Ivan V. Smirnov, Timothy Baldwin, Artem Shelmanov |
| 2025 | NAACL | Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability. | Haonan Li, Xudong Han, Zenan Zhai, Honglin Mu, Hao Wang, Zhenxuan Zhang, Yilin Geng, Shom Lin, Renxi Wang, Artem Shelmanov, Xiangyu Qi, Yuxia Wang, Donghai Hong, Youliang Yuan, Meng Chen, Haoqin Tu, Fajri Koto, Cong Zeng, Tatsuki Kuribayashi, Rishabh Bhardwaj, Bingchen Zhao, Yawen Duan, Yi Liu, Emad A. Alghamdi, Yaodong Yang, Yinpeng Dong, Soujanya Poria, Pengfei Liu, Zhengzhong Liu, Xuguang Ren, Eduard H. Hovy, Iryna Gurevych, Preslav Nakov, Monojit Choudhury, Timothy Baldwin |
| 2025 | NAACL | Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models. | Artem Vazhentsev, Lyudmila Rvanova, Ivan Lazichny, Alexander Panchenko, Maxim Panov, Timothy Baldwin, Artem Shelmanov |
| 2025 | NAACL | NAT: Enhancing Agent Tuning with Negative Samples. | Renxi Wang, Xudong Han, Yixuan Zhang, Timothy Baldwin, Haonan Li |
| 2025 | NAACL | Evaluating Evidence Attribution in Generated Fact Checking Explanations. | Rui Xing, Timothy Baldwin, Jey Han Lau |
| 2024 | ACL | CMMLU: Measuring massive multitask language understanding in Chinese. | Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, Timothy Baldwin |
| 2024 | ACL | Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification. | Ekaterina Fadeeva, Aleksandr Rubashevskii, Artem Shelmanov, Sergey Petrakov, Haonan Li, Hamdy Mubarak, Evgenii Tsymbalov, Gleb Kuzmin, Alexander Panchenko, Timothy Baldwin, Preslav Nakov, Maxim Panov |
| 2024 | ACL | ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic. | Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, Timothy Baldwin |
| 2024 | ACL | Emergent Word Order Universals from Cognitively-Motivated Language Models. | Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin |
| 2024 | ACL | Demystifying Instruction Mixing for Fine-tuning Large Language Models. | Renxi Wang, Haonan Li, Minghao Wu, Yuxia Wang, Xudong Han, Chiyu Zhang, Timothy Baldwin |
| 2024 | ACL | A Chinese Dataset for Evaluating the Safeguards in Large Language Models. | Yuxia Wang, Zenan Zhai, Haonan Li, Xudong Han, Shom Lin, Zhenxuan Zhang, Angela Zhao, Preslav Nakov, Timothy Baldwin |
| 2024 | EACL | Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon. | Fajri Koto, Tilman Beck, Zeerak Talat, Iryna Gurevych, Timothy Baldwin |
| 2024 | EACL | Do-Not-Answer: Evaluating Safeguards in LLMs. | Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, Timothy Baldwin |
| 2024 | EMNLP | BiMediX: Bilingual Medical Mixture of Experts LLM. | Sara Pieri, Sahal Shaji Mullappilly, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman H. Khan, Timothy Baldwin, Hisham Cholakkal |
| 2024 | NAACL | Psychometric Predictive Power of Large Language Models. | Tatsuki Kuribayashi, Yohei Oseki, Timothy Baldwin |
| 2024 | NAACL | Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings. | Chen Liu, Fajri Koto, Timothy Baldwin, Iryna Gurevych |
| 2024 | NAACL | Revisiting subword tokenization: A case study on affixal negation in large language models. | Thinh Truong, Yulia Otmakhova, Karin Verspoor, Trevor Cohn, Timothy Baldwin |
| 2023 | ACL | Unsupervised Paraphrasing of Multiword Expressions. | Takashi Wada, Yuji Matsumoto, Timothy Baldwin, Jey Han Lau |
| 2023 | ACL | NusaCrowd: Open Source Initiative for Indonesian NLP Resources. | Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata, Bryan Wilie, Fajri Koto, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Jennifer Santoso, David Moeljadi, Cahya Wirawan, Frederikus Hudi, Muhammad Satrio Wicaksono, Ivan Halim Parmonangan, Ika Alfina, Ilham Firdausi Putra, Samsul Rahmadani, Yulianti Oenang, Ali Akbar Septiandri, James Jaya, Kaustubh D. Dhole, Arie Ardiyanti Suryani, Rifki Afina Putri, Dan Su, Keith Stevens, Made Nindyatama Nityasya, Muhammad Farid Adilazuarda, Ryan Hadiwijaya, Ryandito Diandaru, Tiezheng Yu, Vito Ghifari, Wenliang Dai, Yan Xu, Dyah Damapuspita, Haryo Akbarianto Wibowo, Cuk Tho, Ichwanul Muslim Karo Karo, Tirana Fatyanosa, Ziwei Ji, Graham Neubig, Timothy Baldwin, Sebastian Ruder, Pascale Fung, Herry Sujaini, Sakriani Sakti, Ayu Purwarianti |
| 2023 | ACL | Cost-effective Distillation of Large Language Models. | Sayantan Dasgupta, Trevor Cohn, Timothy Baldwin |
| 2023 | EACL | Fair Enough: Standardizing Evaluation and Model Selection for Fairness Research in NLP. | Xudong Han, Timothy Baldwin, Trevor Cohn |
| 2023 | EMNLP | Unsupervised Lexical Simplification with Context Augmentation. | Takashi Wada, Timothy Baldwin, Jey Han Lau |
| 2023 | EMNLP | LM-Polygraph: Uncertainty Estimation for Language Models. | Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, Timothy Baldwin, Artem Shelmanov |
| 2023 | EMNLP | More than Votes? Voting and Language based Partisanship in the US Supreme Court. | Biaoyan Fang, Trevor Cohn, Timothy Baldwin, Lea Frermann |
| 2023 | EMNLP | Robustness Tests for Automatic Machine Translation Metrics with Adversarial Attacks. | Yichen Huang, Timothy Baldwin |
| 2023 | EMNLP | Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLU. | Fajri Koto, Nurul Aisyah, Haonan Li, Timothy Baldwin |
| 2023 | ICLR | Everybody Needs Good Neighbours: An Unsupervised Locality-based Method for Bias Mitigation. | Xudong Han, Timothy Baldwin, Trevor Cohn |
| 2023 | IJCNLP | It's not only What You Say, It's also Who It's Said to: Counterfactual Analysis of Interactive Behavior in the Courtroom. | Biaoyan Fang, Trevor Cohn, Timothy Baldwin, Lea Frermann |
| 2023 | IJCNLP | Uncertainty Estimation for Debiased Models: Does Fairness Hurt Reliability? | Gleb Kuzmin, Artem Vazhentsev, Artem Shelmanov, Xudong Han, Simon Suster, Maxim Panov, Alexander Panchenko, Timothy Baldwin |
| 2023 | IJCNLP | Location Aware Modular Biencoder for Tourism Question Answering. | Haonan Li, Martin Tomko, Timothy Baldwin |
| 2022 | ACE | Online Examinations in a Large Australian CS1 Course. | Bryn Jeffries, Timothy Baldwin, Marion Zalk |
| 2022 | ACL | The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature. | Yulia Otmakhova, Karin Verspoor, Timothy Baldwin, Jey Han Lau |
| 2022 | ACL | One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia. | Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder |
| 2022 | ACL | What does it take to bake a cake? The RecipeRef corpus and anaphora resolution in procedural text. | Biaoyan Fang, Timothy Baldwin, Karin Verspoor |
| 2022 | COLING | Unsupervised Lexical Substitution with Decontextualised Embeddings. | Takashi Wada, Timothy Baldwin, Yuji Matsumoto, Jey Han Lau |
| 2022 | COLING | LED down the rabbit hole: exploring the potential of global attention for biomedical multi-document summarisation. | Yulia Otmakhova, Thinh Hung Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor, Jey Han Lau |
| 2022 | COLING | LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization. | Fajri Koto, Timothy Baldwin, Jey Han Lau |
| 2022 | COLING | Noisy Label Regularisation for Textual Regression. | Yuxia Wang, Timothy Baldwin, Karin Verspoor |
| 2022 | ECIR | The ChEMU 2022 Evaluation Campaign: Information Extraction in Chemical Patents. | Yuan Li, Biaoyan Fang, Jiayuan He, Hiyori Yoshikawa, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zenan Zhai, Zubair Afzal, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
| 2022 | EMNLP | M3: Multi-level dataset for Multi-document summarisation of Medical studies. | Yulia Otmakhova, Karin Verspoor, Timothy Baldwin, Antonio Jimeno-Yepes, Jey Han Lau |
| 2022 | EMNLP | Balancing out Bias: Achieving Fairness Through Balanced Training. | Xudong Han, Timothy Baldwin, Trevor Cohn |
| 2022 | EMNLP | FairLib: A Unified Framework for Assessing and Improving Fairness. | Xudong Han, Aili Shen, Yitong Li, Lea Frermann, Timothy Baldwin, Trevor Cohn |
| 2022 | IJCNLP | Systematic Evaluation of Predictive Fairness. | Xudong Han, Aili Shen, Trevor Cohn, Timothy Baldwin, Lea Frermann |
| 2022 | IJCNLP | Does Representational Fairness Imply Empirical Fairness? | Aili Shen, Xudong Han, Trevor Cohn, Timothy Baldwin, Lea Frermann |
| 2022 | IJCNLP | Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation. | Thinh Hung Truong, Yulia Otmakhova, Timothy Baldwin, Trevor Cohn, Jey Han Lau, Karin Verspoor |
| 2022 | NAACL | CULG: Commercial Universal Language Generation. | Haonan Li, Yameng Huang, Yeyun Gong, Jian Jiao, Ruofei Zhang, Timothy Baldwin, Nan Duan |
| 2022 | NAACL | MultiSpanQA: A Dataset for Multi-Span Question Answering. | Haonan Li, Martin Tomko, Maria Vasardani, Timothy Baldwin |
| 2022 | NAACL | Optimising Equal Opportunity Fairness in Model Training. | Aili Shen, Xudong Han, Trevor Cohn, Timothy Baldwin, Lea Frermann |
| 2022 | NAACL | Improving negation detection with negation-focused pre-training. | Thinh Hung Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor |
| 2021 | ACL | Decoupling Adversarial Training for Fair NLP. | Xudong Han, Timothy Baldwin, Trevor Cohn |
| 2021 | ACL | Evaluating the Efficacy of Summarization Evaluation across Languages. | Fajri Koto, Jey Han Lau, Timothy Baldwin |
| 2021 | EACL | ChEMU-Ref: A Corpus for Modeling Anaphora Resolution in the Chemical Domain. | Biaoyan Fang, Christian Druckenbrodt, Saber A. Akhondi, Jiayuan He, Timothy Baldwin, Karin Verspoor |
| 2021 | EACL | Diverse Adversaries for Mitigating Bias in Training. | Xudong Han, Timothy Baldwin, Trevor Cohn |
| 2021 | EACL | Top-down Discourse Parsing via Sequence Labelling. | Fajri Koto, Jey Han Lau, Timothy Baldwin |
| 2021 | EACL | On the (In)Effectiveness of Images for Text Classification. | Chunpeng Ma, Aili Shen, Hiyori Yoshikawa, Tomoya Iwakura, Daniel Beck, Timothy Baldwin |
| 2021 | ECIR | ChEMU 2021: Reaction Reference Resolution and Anaphora Resolution in Chemical Patents. | Jiayuan He, Biaoyan Fang, Hiyori Yoshikawa, Yuan Li, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zubair Afzal, Zenan Zhai, Lawrence Cavedon, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
| 2021 | ECIR | Brief Description of COVID-SEE: The Scientific Evidence Explorer for COVID-19 Related Research. | Karin Verspoor, Simon Suster, Yulia Otmakhova, Shevon Mendis, Zenan Zhai, Biaoyan Fang, Jey Han Lau, Timothy Baldwin, Antonio Jimeno-Yepes, David Martnez |
| 2021 | EMNLP | KFCNet: Knowledge Filtering and Contrastive Learning for Generative Commonsense Reasoning. | Haonan Li, Yeyun Gong, Jian Jiao, Ruofei Zhang, Timothy Baldwin, Nan Duan |
| 2021 | EMNLP | IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization. | Fajri Koto, Jey Han Lau, Timothy Baldwin |
| 2021 | EMNLP | 'Just What do You Think You're Doing, Dave?' A Checklist for Responsible Data Use in NLP. | Anna Rogers, Timothy Baldwin, Kobi Leins |
| 2021 | EMNLP | Fairness-aware Class Imbalanced Learning. | Shivashankar Subramanian, Afshin Rahimi, Timothy Baldwin, Trevor Cohn, Lea Frermann |
| 2021 | EMNLP | Evaluating Debiasing Techniques for Intersectional Biases. | Shivashankar Subramanian, Xudong Han, Timothy Baldwin, Trevor Cohn, Lea Frermann |
| 2021 | NAACL | Automatic Classification of Neutralization Techniques in the Narrative of Climate Change Scepticism. | Shraey Bhatia, Jey Han Lau, Timothy Baldwin |
| 2021 | NAACL | Discourse Probing of Pretrained Language Models. | Fajri Koto, Jey Han Lau, Timothy Baldwin |
| 2021 | SIGdial | A Simple yet Effective Method for Sentence Ordering. | Aili Shen, Timothy Baldwin |
| 2020 | ACE | Online Tutoring to Support Programming Exercises. | Bryn Jeffries, Timothy Baldwin, Marion Zalk, Ben Taylor |
| 2020 | ACL | Give Me Convenience and Give Her Death: Who Should Decide What Uses of NLP are Appropriate, and on What Basis? | Kobi Leins, Jey Han Lau, Timothy Baldwin |
| 2020 | ACL | Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. | Nitika Mathur, Timothy Baldwin, Trevor Cohn |
| 2020 | COLING | IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. | Fajri Koto, Afshin Rahimi, Jey Han Lau, Timothy Baldwin |
| 2020 | COLING | Target Word Masking for Location Metonymy Resolution. | Haonan Li, Maria Vasardani, Martin Tomko, Timothy Baldwin |
| 2020 | COLING | WikiUMLS: Aligning UMLS to Wikipedia via Cross-lingual Neural Ranking. | Afshin Rahimi, Timothy Baldwin, Karin Verspoor |
| 2020 | ECIR | ChEMU: Named Entity Recognition and Event Extraction of Chemical Reactions from Patents. | Dat Quoc Nguyen, Zenan Zhai, Hiyori Yoshikawa, Biaoyan Fang, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Saber A. Akhondi, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
| 2020 | EMNLP | Improved Topic Representations of Medical Documents to Assist COVID-19 Literature Exploration. | Yulia Otmakhova, Karin Verspoor, Timothy Baldwin, Simon Suster |
| 2020 | IJCNLP | Liputan6: A Large-scale Indonesian Dataset for Text Summarization. | Fajri Koto, Jey Han Lau, Timothy Baldwin |
| 2019 | ACL | Semi-supervised Stochastic Multi-Domain Learning using Variational Inference. | Yitong Li, Timothy Baldwin, Trevor Cohn |
| 2019 | ACL | Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation. | Nitika Mathur, Timothy Baldwin, Trevor Cohn |
| 2019 | ADCS | Differences in language use: Insights from job and talent search. | Bahar Salehi, Borhan Kazimipour, Timothy Baldwin |
| 2019 | EMNLP | Deep Ordinal Regression for Pledge Specificity Prediction. | Shivashankar Subramanian, Trevor Cohn, Timothy Baldwin |
| 2019 | NAACL | Contextualization of Morphological Inflection. | Ekaterina Vylomova, Ryan Cotterell, Trevor Cohn, Timothy Baldwin, Jason Eisner |
| 2018 | ACL | Deep-speare: A joint neural model of poetic language, meter and rhyme. | Jey Han Lau, Trevor Cohn, Timothy Baldwin, Julian Brooke, Adam Hammond |
| 2018 | ACL | Semi-supervised User Geolocation via Graph Convolutional Networks. | Afshin Rahimi, Trevor Cohn, Timothy Baldwin |
| 2018 | ACL | Towards Robust and Privacy-preserving Text Representations. | Yitong Li, Timothy Baldwin, Trevor Cohn |
| 2018 | ACL | Narrative Modeling with Memory Chains and Semantic Supervision. | Fei Liu, Trevor Cohn, Timothy Baldwin |
| 2018 | ACL | Content-based Popularity Prediction of Online Petitions Using a Deep Regression Model. | Shivashankar Subramanian, Timothy Baldwin, Trevor Cohn |
| 2018 | COLING | Encoding Sentiment Information into Word Vectors for Sentiment Analysis. | Zhe Ye, Fang Li, Timothy Baldwin |
| 2018 | EMNLP | Topic Intrusion for Automatic Topic Model Evaluation. | Shraey Bhatia, Jey Han Lau, Timothy Baldwin |
| 2018 | ICTIR | Multitask Learning for Query Segmentation in Job Search. | Bahar Salehi, Fei Liu, Timothy Baldwin, Wilson Wong |
| 2018 | ICWSM | Detecting Misflagged Duplicate Questions in Community Question-Answering Archives. | Doris Hoogeveen, Andrew Bennett, Yitong Li, Karin M. Verspoor, Timothy Baldwin |
| 2018 | NAACL | What's in a Domain? Learning Domain-Robust Text Representations using Adversarial Training. | Yitong Li, Timothy Baldwin, Trevor Cohn |
| 2018 | NAACL | Recurrent Entity Networks with Delayed Memory Update for Targeted Aspect-Based Sentiment Analysis. | Fei Liu, Trevor Cohn, Timothy Baldwin |
| 2018 | NAACL | Hierarchical Structured Model for Fine-to-Coarse Manifesto Text Analysis. | Shivashankar Subramanian, Trevor Cohn, Timothy Baldwin |
| 2018 | SIGIR | A Living Lab Study of Query Amendment in Job Search. | Bahar Salehi, Damiano Spina, Alistair Moffat, Sargol Sadeghi, Falk Scholer, Timothy Baldwin, Lawrence Cavedon, Mark Sanderson, Wilson Wong, Justin Zobel |
| 2017 | ACL | Topically Driven Neural Language Model. | Jey Han Lau, Timothy Baldwin, Trevor Cohn |
| 2017 | ACL | A Neural Model for User Geolocation and Lexical Dialectology. | Afshin Rahimi, Trevor Cohn, Timothy Baldwin |
| 2017 | CoNLL | An Automatic Approach for Document-level Topic Model Evaluation. | Shraey Bhatia, Jey Han Lau, Timothy Baldwin |
| 2017 | EACL | Robust Training under Linguistic Adversity. | Yitong Li, Trevor Cohn, Timothy Baldwin |
| 2017 | EACL | Context-Aware Prediction of Derivational Word-forms. | Ekaterina Vylomova, Ryan Cotterell, Timothy Baldwin, Trevor Cohn |
| 2017 | EACL | Multimodal Topic Labelling. | Ionut Sorodoc, Jey Han Lau, Nikolaos Aletras, Timothy Baldwin |
| 2017 | EACL | Improving Evaluation of Document-level Machine Translation Quality Estimation. | Yvette Graham, Qingsong Ma, Timothy Baldwin, Qun Liu, Carla Parra Escartn, Carolina Scarton |
| 2017 | EMNLP | Further Investigation into Reference Bias in Monolingual Evaluation of Machine Translation. | Qingsong Ma, Yvette Graham, Timothy Baldwin, Qun Liu |
| 2017 | EMNLP | Sequence Effects in Crowdsourced Annotations. | Nitika Mathur, Timothy Baldwin, Trevor Cohn |
| 2017 | EMNLP | Sub-character Neural Language Modelling in Japanese. | Viet Nguyen, Julian Brooke, Timothy Baldwin |
| 2017 | EMNLP | Continuous Representation of Location for Geolocation and Lexical Dialectology using Mixture Density Networks. | Afshin Rahimi, Timothy Baldwin, Trevor Cohn |
| 2017 | IJCNLP | Capturing Long-range Contextual Dependencies with Memory-enhanced Conditional Random Fields. | Fei Liu, Timothy Baldwin, Trevor Cohn |
| 2017 | WWW | Pairwise Webpage Coreference Classification Using Distant Supervision. | Shivashankar Subramanian, Timothy Baldwin, Julian Brooke, Trevor Cohn |
| 2017 | SIGIR | Understanding User Behavior in Job and Talent Search: An Initial Investigation. | Damiano Spina, Maria Maistro, Yongli Ren, Sargol Sadeghi, Wilson Wong, Timothy Baldwin, Lawrence Cavedon, Alistair Moffat, Mark Sanderson, Falk Scholer, Justin Zobel |
| 2016 | ACL | LexSemTm: A Semantic Dataset Based on All-words Unsupervised Sense Distribution Learning. | Andrew Bennett, Timothy Baldwin, Jey Han Lau, Diana McCarthy, Francis Bond |
| 2016 | ACL | Bootstrapped Text-level Named Entity Recognition for Literature. | Julian Brooke, Adam Hammond, Timothy Baldwin |
| 2016 | ACL | pigeo: A Python Geotagging Tool. | Afshin Rahimi, Trevor Cohn, Timothy Baldwin |
| 2016 | ACL | Take and Took, Gaggle and Goose, Book and Read: Evaluating the Utility of Vector Differences for Lexical Relation Learning. | Ekaterina Vylomova, Laura Rimell, Trevor Cohn, Timothy Baldwin |
| 2016 | COLING | Automatic Labelling of Topics with Neural Embeddings. | Shraey Bhatia, Jey Han Lau, Timothy Baldwin |
| 2016 | COLING | Is all that Glitters in Machine Translation Quality Estimation really Gold? | Yvette Graham, Timothy Baldwin, Meghan Dowling, Maria Eskevich, Teresa Lynn, Lamia Tounsi |
| 2016 | COLING | Determining the Multiword Expression Inventory of a Surprise Language. | Bahar Salehi, Paul Cook, Timothy Baldwin |
| 2016 | EMNLP | Learning Robust Representations of Text. | Yitong Li, Trevor Cohn, Timothy Baldwin |
| 2016 | EMNLP | Named Entity Recognition for Novel Types by Transfer Learning. | Lizhen Qu, Gabriela Ferraro, Liyuan Zhou, Weiwei Hou, Timothy Baldwin |
| 2016 | LREC | Evaluating a Topic Modelling Approach to Measuring Corpus Similarity. | Richard Fothergill, Paul Cook, Timothy Baldwin |
| 2016 | NAACL | The Sensitivity of Topic Coherence Evaluation to Topic Cardinality. | Jey Han Lau, Timothy Baldwin |
| 2016 | SIGIR | Quit While Ahead: Evaluating Truncated Rankings. | Fei Liu, Alistair Moffat, Timothy Baldwin, Xiuzhen Zhang |
| 2015 | ACL | Twitter User Geolocation Using a Unified Text and Network Prediction Model. | Afshin Rahimi, Trevor Cohn, Timothy Baldwin |
| 2015 | ADCS | CQADupStack: A Benchmark Data Set for Community Question-Answering Research. | Doris Hoogeveen, Karin M. Verspoor, Timothy Baldwin |
| 2015 | CIKM | TM 2015 - Topic Models: Post-Processing and Applications Workshop. | Nikolaos Aletras, Jey Han Lau, Timothy Baldwin, Mark Stevenson |
| 2015 | CIKM | A Probabilistic Rating Auto-encoder for Personalized Recommender Systems. | Huizhi Liang, Timothy Baldwin |
| 2015 | CoNLL | Big Data Small Data, In Domain Out-of Domain, Known Word Unknown Word: The Impact of Word Representations on Sequence Labelling Tasks. | Lizhen Qu, Gabriela Ferraro, Liyuan Zhou, Weiwei Hou, Nathan Schneider, Timothy Baldwin |
| 2015 | NAACL | Accurate Evaluation of Segment-level Machine Translation Metrics. | Yvette Graham, Timothy Baldwin, Nitika Mathur |
| 2015 | NAACL | Exploiting Text and Network Context for Geolocation of Social Media Users. | Afshin Rahimi, Duy Vu, Trevor Cohn, Timothy Baldwin |
| 2015 | NAACL | A Word Embedding Approach to Predicting the Compositionality of Multiword Expressions. | Bahar Salehi, Paul Cook, Timothy Baldwin |
| 2014 | ACL | Automatic Detection of Multilingual Dictionaries on the Web. | Gintare Grigonyte, Timothy Baldwin |
| 2014 | ACL | Learning Word Sense Distributions, Detecting Unattested Senses and Identifying Novel Senses Using Topic Models. | Jey Han Lau, Paul Cook, Diana McCarthy, Spandana Gella, Timothy Baldwin |
| 2014 | COLING | Novel Word-sense Identification. | Paul Cook, Jey Han Lau, Diana McCarthy, Timothy Baldwin |
| 2014 | EACL | One Sense per Tweeter ... and Other Lexical Semantic Tales of Twitter. | Spandana Gella, Paul Cook, Timothy Baldwin |
| 2014 | EACL | Is Machine Translation Getting Better over Time? | Yvette Graham, Timothy Baldwin, Alistair Moffat, Justin Zobel |
| 2014 | EACL | Machine Reading Tea Leaves: Automatically Evaluating Topic Coherence and Topic Model Quality. | Jey Han Lau, David Newman, Timothy Baldwin |
| 2014 | EACL | Using Distributional Similarity of Multi-way Translations to Predict Multiword Expression Compositionality. | Bahar Salehi, Paul Cook, Timothy Baldwin |
| 2014 | EMNLP | Testing for Significance of Increased Correlation with Human Judgment. | Yvette Graham, Timothy Baldwin |
| 2014 | EMNLP | Detecting Non-compositional MWE Components using Wiktionary. | Bahar Salehi, Paul Cook, Timothy Baldwin |
| 2013 | ACL | A Stacking-based Approach to Twitter User Geolocation Prediction. | Bo Han, Paul Cook, Timothy Baldwin |
| 2013 | IJCNLP | How Noisy Social Media Text, How Diffrnt Social Media Sources? | Timothy Baldwin, Paul Cook, Marco Lui, Andrew MacKinlay, Li Wang |
| 2013 | IJCNLP | Unsupervised Word Class Induction for Under-resourced Languages: A Case Study on Indonesian. | Meladel Mistica, Jey Han Lau, Timothy Baldwin |
| 2012 | ACL | langid.py: An Off-the-shelf Language Identification Tool. | Marco Lui, Timothy Baldwin |
| 2012 | COLING | Geolocation Prediction in Social Media Data by Finding Location Indicative Words. | Bo Han, Paul Cook, Timothy Baldwin |
| 2012 | COLING | On-line Trend Analysis with Topic Models: \#twitter Trends Detection Topic Model Online. | Jey Han Lau, Nigel Collier, Timothy Baldwin |
| 2012 | COLING | Bayesian Text Segmentation for Index Term Identification and Keyphrase Extraction. | David Newman, Nagendra Koilada, Jey Han Lau, Timothy Baldwin |
| 2012 | COLING | The Utility of Discourse Structure in Identifying Resolved Threads in Technical User Forums. | Li Wang, Su Nam Kim, Timothy Baldwin |
| 2012 | EACL | A Support Platform for Event Detection using Social Intelligence. | Timothy Baldwin, Paul Cook, Bo Han, Aaron Harwood, Shanika Karunasekera, Masud Moshtaghi |
| 2012 | EACL | Word Sense Induction for Novel Sense Detection. | Jey Han Lau, Paul Cook, Diana McCarthy, David Newman, Timothy Baldwin |
| 2012 | EMNLP | Automatically Constructing a Normalisation Dictionary for Microblogs. | Bo Han, Paul Cook, Timothy Baldwin |
| 2012 | NAACL | Evaluating a Morphological Analyser of Inuktitut. | Jeremy Nicholson, Trevor Cohn, Timothy Baldwin |
| 2012 | PACLIC | Social Media: Friend or Foe of Natural Language Processing? | Timothy Baldwin |
| 2012 | PACLIC | Extracting Keywords from Multi-party Live Chats. | Su Nam Kim, Timothy Baldwin |
| 2012 | PACLIC | Classifying Dialogue Acts in Multi-party Live Chats. | Su Nam Kim, Lawrence Cavedon, Timothy Baldwin |
| 2012 | PACLIC | Deep Lexical Acquisition of Type Properties in Low-resource Languages: A Case Study in Wambaya. | Jeremy Nicholson, Rachel Nordlinger, Timothy Baldwin |
| 2011 | ACL | Collective Classification of Congressional Floor-Debate Transcripts. | Clinton Burfoot, Steven Bird, Timothy Baldwin |
| 2011 | ACL | Lexical Normalisation of Short Text Messages: Makn Sens a #twitter. | Bo Han, Timothy Baldwin |
| 2011 | ACL | Automatic Labelling of Topic Models. | Jey Han Lau, Karl Grieser, David Newman, Timothy Baldwin |
| 2011 | ACL | Relation Guided Bootstrapping of Semantic Lexicons. | Tara McIntosh, Lars Yencken, James R. Curran, Timothy Baldwin |
| 2011 | EMNLP | Predicting Thread Discourse Structure over Technical Web Forums. | Li Wang, Marco Lui, Su Nam Kim, Joakim Nivre, Timothy Baldwin |
| 2011 | IJCAI | Modelling and Predicting Movements of Museum Visitors: A Simulation Framework for Assessing the Impact of Sensor Noise on Model Performance. | Fabian Bohnert, Ingrid Zukerman, David W. Albrecht, Timothy Baldwin |
| 2011 | IJCNLP | Fleshing it out: A Supervised Approach to MWE-token and MWE-type Classification. | Richard Fothergill, Timothy Baldwin |
| 2011 | IJCNLP | Cross-domain Feature Selection for Language Identification. | Marco Lui, Timothy Baldwin |
| 2011 | IJCNLP | Treeblazing: Using External Treebanks to Filter Parse Forests for Parse Selection and Treebanking. | Andrew MacKinlay, Rebecca Dridan, Dan Flickinger, Stephan Oepen, Timothy Baldwin |
| 2011 | IUI | Predicting and compensating for lexicon access errors. | Lars Yencken, Timothy Baldwin |
| 2011 | PACLIC | In Situ Text Summarisation for Museum Visitors. | Timothy Baldwin, Patrick Ye, Fabian Bohnert, Ingrid Zukerman |
| 2011 | PACLIC | Word classes in Indonesian: A linguistic reality or a convenient fallacy in natural language processing? | Meladel Mistica, Timothy Baldwin, I Wayan Arka |
| 2010 | COLING | PanLex and LEXTRACT: Translating all Words of all Languages of the World. | Timothy Baldwin, Jonathan Pool, Susan M. Colowick |
| 2010 | COLING | Evaluating N-gram based Evaluation Metrics for Automatic Keyphrase Extraction. | Su Nam Kim, Timothy Baldwin, Min-Yen Kan |
| 2010 | COLING | Best Topic Word Selection for Topic Labelling. | Jey Han Lau, David Newman, Sarvnaz Karimi, Timothy Baldwin |
| 2010 | CoNLL | Tagging and Linking Web Forum Posts. | Su Nam Kim, Li Wang, Timothy Baldwin |
| 2010 | EMNLP | Unsupervised Parse Selection for HPSG. | Rebecca Dridan, Timothy Baldwin |
| 2010 | EMNLP | Classifying Dialogue Acts in One-on-One Live Chats. | Su Nam Kim, Lawrence Cavedon, Timothy Baldwin |
| 2010 | NAACL | Language Identification: The Long and the Short of the Matter. | Timothy Baldwin, Marco Lui |
| 2010 | NAACL | Automatic Evaluation of Topic Coherence. | David Newman, Jey Han Lau, Karl Grieser, Timothy Baldwin |
| 2010 | NAACL | Chart Mining-based Lexical Acquisition with Precision Grammars. | Yi Zhang, Timothy Baldwin, Valia Kordoni, David Martnez, Jeremy Nicholson |
| 2009 | ACL | Automatic Satire Detection: Are You Having a Laugh? | Clinton Burfoot, Timothy Baldwin |
| 2009 | CIKM | Experiments on pattern-based relation learning. | Willy Yap, Timothy Baldwin |
| 2009 | NAACL | Recognising the Predicate-argument Structure of Tagalog. | Meladel Mistica, Timothy Baldwin |
| 2009 | NAACL | Web and Corpus Methods for Malay Count Classifier Prediction. | Jeremy Nicholson, Timothy Baldwin |
| 2008 | AAAI | Towards Automatic Animated Storyboarding. | Patrick Ye, Timothy Baldwin |
| 2008 | ACL | Improving Parsing and PP Attachment Performance with Sense Information. | Eneko Agirre, Timothy Baldwin, David Martnez |
| 2008 | COLING | Applying Discourse Analysis and Data Mining Methods to Spoken OSCE Assessments. | Meladel Mistica, Timothy Baldwin, Marisa Cordella, Simon Musgrave |
| 2008 | COLING | Measuring and Predicting Orthographic Associations: Modelling the Similarity of Japanese Kanji. | Lars Yencken, Timothy Baldwin |
| 2008 | ECAI | Orthographic similarity search for dictionary lookup of Japanese words. | Lars Yencken, Timothy Baldwin |
| 2008 | IJCNLP | MRD-based Word Sense Disambiguation: Further Extending Lesk. | Timothy Baldwin, Su Nam Kim, Francis Bond, Sanae Fujita, David Martnez, Takaaki Tanaka |
| 2008 | IJCNLP | Benchmarking Noun Compound Interpretation. | Su Nam Kim, Timothy Baldwin |
| 2008 | LREC | Evaluating and Extending the Coverage of HPSG Grammars: A Case Study for German. | Jeremy Nicholson, Valia Kordoni, Yi Zhang, Timothy Baldwin, Rebecca Dridan |
| 2007 | AAAI | Disambiguating Noun Compounds. | Su Nam Kim, Timothy Baldwin |
| 2007 | EMNLP | Word Sense Disambiguation Incorporating Lexical and Structural Semantic Information. | Takaaki Tanaka, Francis Bond, Timothy Baldwin, Sanae Fujita, Chikara Hashimoto |
| 2007 | PACLIC | Scalable Deep Linguistic Processing: Mind the Lexical Gap. | Timothy Baldwin |
| 2006 | ACL | Interpreting Semantic Relations in Noun Compounds via Verb Semantics. | Su Nam Kim, Timothy Baldwin |
| 2006 | EMNLP | Multilingual Deep Lexical Acquisition for HPSGs via Supertagging. | Phil Blunsom, Timothy Baldwin |
| 2006 | LREC | Open Source Corpus Analysis Tools for Malay. | Timothy Baldwin, Su'ad Awab |
| 2006 | LREC | Reconsidering Language Identification for Written Language Resources. | Baden Hughes, Timothy Baldwin, Steven Bird, Jeremy Nicholson, Andrew MacKinlay |
| 2005 | IJCNLP | Automatic Interpretation of Noun Compounds Using WordNet Similarity. | Su Nam Kim, Timothy Baldwin |
| 2005 | IJCNLP | Semantic Role Labelling of Prepositional Phrases. | Patrick Ye, Timothy Baldwin |
| 2004 | LREC | Road-testing the English Resource Grammar Over the British National Corpus. | Timothy Baldwin, Emily M. Bender, Dan Flickinger, Ara Kim, Stephan Oepen |
| 2004 | LREC | Evaluating the FOKS Error Model. | Slaven Bilac, Timothy Baldwin, Hozumi Tanaka |
| 2004 | LREC | A Multilingual Database of Idioms. | Aline Villavicencio, Timothy Baldwin, Benjamin Waldron |
| 2004 | PACLIC | Automatic Discovery of Telic and Agentive Roles from Corpus Data. | Ichiro Yamada, Timothy Baldwin |
| 2003 | ACL | Learning the Countability of English Nouns from Corpus Data. | Timothy Baldwin, Francis Bond |
| 2003 | EMNLP | A Plethora of Methods for Learning English Countability. | Timothy Baldwin, Francis Bond |
| 2002 | CICLING | Multiword Expressions: A Pain in the Neck for NLP. | Ivan A. Sag, Timothy Baldwin, Francis Bond, Ann A. Copestake, Dan Flickinger |
| 2002 | COLING | Bringing the Dictionary to the User: The FOKS System. | Slaven Bilac, Timothy Baldwin, Hozumi Tanaka |
| 2002 | CoNLL | Extracting the Unextractable: A Case Study on Verb-particles. | Timothy Baldwin, Aline Villavicencio |
| 2002 | LREC | Enhanced Japanese Electronic Dictionary Look-up. | Timothy Baldwin, Slaven Bilac, Ryo Okumura, Takenobu Tokunaga, Hozumi Tanaka |
| 2002 | LREC | Multiword expressions: linguistic precision and reusability. | Ann A. Copestake, Fabre Lambeau, Aline Villavicencio, Francis Bond, Timothy Baldwin, Ivan A. Sag, Dan Flickinger |
| 2001 | ACL | Low-cost, High-Performance Translation Retrieval: Dumber is Better. | Timothy Baldwin |
| 2000 | COLING | The Effects of Word Order and Segmentation on Translation Retrieval Performance. | Timothy Baldwin, Hozumi Tanaka |
| 2000 | PACLIC | Verb Alternations and Japanese : How, What and Where. | Timothy Baldwin, Hozumi Tanaka |