Skip to content

Bilal Piot

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

27

Venues

6

Active years

2014–2025

Best venue rank

A*

Where they publish

Papers

27 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRRRM: Robust Reward Model Training Mitigates Reward Hacking.Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Zhe Liu, Yuan Liu, Bilal Piot, Abe Ittycheriah, Aviral Kumar, Mohammad Saleh
2025ICLRBuilding Math Agents with Multi-Turn Iterative Preference Learning.Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg, Zhen Qin, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, Chi Jin, Tong Zhang, Tianqi Liu
2025ICLRLearning from negative feedback, or positive feedback or both.Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari, Jost Tobias Springenberg, Tim Hertweck, Michael Bloesch, Rishabh Joshi, Thomas Lampe, Junhyuk Oh, Nicolas Heess, Jonas Buchli, Martin A. Riedmiller
2024AISTATSA General Theoretical Paradigm to Understand Learning from Human Preferences.Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello
2024ICLRUnlocking the Power of Representations in Long-term Novelty-based Exploration.Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Pablo Sprechmann, Leopoldo Sarra, Oliver Groth, Michal Valko, Bilal Piot
2024ICMLHuman Alignment of Large Language Models through Online Preference Optimisation.Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot
2024ICMLNash Learning from Human Feedback.Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot
2024ICMLGeneralized Preference Optimization: A Unified Approach to Offline Alignment.Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot
2023ICMLThe Edge of Orthogonality: A Simple View of What Makes BYOL Tick.Pierre Harvey Richemond, Allison C. Tam, Yunhao Tang, Florian Strub, Bilal Piot, Felix Hill
2023ICMLUnderstanding Self-Predictive Learning for Reinforcement Learning.Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
2022ICLREmergent Communication at Scale.Rahma Chaabouni, Florian Strub, Florent Altch, Eugene Tarassov, Corentin Tallec, Elnaz Davoodi, Kory Wallace Mathewson, Olivier Tieleman, Angeliki Lazaridou, Bilal Piot
2020ICLRNever Give Up: Learning Directed Exploration Strategies.Adri Puigdomnech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martn Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell
2020ICMLAgent57: Outperforming the Atari Human Benchmark.Adri Puigdomnech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Charles Blundell
2020ICMLBootstrap Latent-Predictive Representations for Multitask Reinforcement Learning.Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar
2018AAAIRainbow: Combining Improvements in Deep Reinforcement Learning.Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, David Silver
2018AAAIDeep Q-learning From Demonstrations.Todd Hester, Matej Vecerk, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, Audrunas Gruslys
2018AISTATSActor-Critic Fictitious Play in Simultaneous Move Multistage Games.Julien Prolat, Bilal Piot, Olivier Pietquin
2018ICLRNoisy Networks For Exploration.Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg
2018ICLRThe Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning.Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos
2017AISTATSLearning Nash Equilibrium for General-Sum Markov Games from Batch Data.Julien Prolat, Florian Strub, Bilal Piot, Olivier Pietquin
2017IJCAIEnd-to-end optimization of goal-driven and visually grounded dialogue systems.Florian Strub, Harm de Vries, Jrmie Mary, Bilal Piot, Aaron C. Courville, Olivier Pietquin
2016AISTATSOn the Use of Non-Stationary Strategies for Solving Two-Player Zero-Sum Markov Games.Julien Prolat, Bilal Piot, Bruno Scherrer, Olivier Pietquin
2016ICMLSoftened Approximate Policy Iteration for Markov Games.Julien Prolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, Olivier Pietquin
2015ICMLApproximate Dynamic Programming for Two-Player Zero-Sum Markov Games.Julien Prolat, Bruno Scherrer, Bilal Piot, Olivier Pietquin
2015ICMLImitation Learning Applied to Embodied Conversational Agents.Bilal Piot, Olivier Pietquin, Matthieu Geist
2015IJCAIInverse Reinforcement Learning in Relational Domains.Thibaut Munzer, Bilal Piot, Matthieu Geist, Olivier Pietquin, Manuel Lopes
2014InterspeechPredicting when to laugh with structured classification.Bilal Piot, Olivier Pietquin, Matthieu Geist