| 2025 | IJCAI | Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture. | John Burden, Marko Tesic, Lorenzo Pacchiardi, Jos Hernndez-Orallo |
| 2024 | AAAI | Your Prompt Is My Command: On Assessing the Human-Centred Generality of Multimodal Models (Abstract Reprint). | Wout Schellaert, Fernando Martnez-Plumed, Karina Vold, John Burden, Pablo A. M. Casares, Bao Sheng Loe, Roi Reichart, Sen higeartaigh, Anna Korhonen, Jos Hernndez-Orallo |
| 2022 | AAAI | Oases of Cooperation: An Empirical Evaluation of Reinforcement Learning in the Iterated Prisoner's Dilemma. | Peter Barnett, John Burden |
| 2022 | AAAI | How General-Purpose Is a Language Model? Usefulness and Safety with Human Prompters in the Wild. | Pablo Antonio Moreno Casares, Bao Sheng Loe, John Burden, Sen higeartaigh, Jos Hernndez-Orallo |
| 2022 | IJCAI | Not a Number: Identifying Instance Features for Capability-Oriented Evaluation. | Ryan Burnell, John Burden, Danaja Rutar, Konstantinos Voudouris, Lucy Cheke, Jos Hernndez-Orallo |
| 2022 | IJCAI | Evaluating Object Permanence in Embodied Agents using the Animal-AI Environment. | Konstantinos Voudouris, Niall Donnelly, Danaja Rutar, Ryan Burnell, John Burden, Jos Hernndez-Orallo, Lucy Cheke |
| 2021 | AAAI | Negative Side Effects and AI Agent Indicators: Experiments in SafeLife. | John Burden, Jos Hernndez-Orallo, Sen higeartaigh |
| 2020 | AAAI | Exploring AI Safety in Degrees: Generality, Capability and Control. | John Burden, Jos Hernndez-Orallo |
| 2020 | ECAI | Uniform State Abstraction for Reinforcement Learning. | John Burden, Daniel Kudenko |