Bilal Piot
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
27
Venues
6
Active years
2014–2025
Best venue rank
A*
Where they publish
Papers
27 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | ICLR | RRM: Robust Reward Model Training Mitigates Reward Hacking. | Tianqi Liu, Wei Xiong, Jie Ren, Lichang Chen, Junru Wu, Rishabh Joshi, Yang Gao, Jiaming Shen, Zhen Qin, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Zhe Liu, Yuan Liu, Bilal Piot, Abe Ittycheriah, Aviral Kumar, Mohammad Saleh |
| 2025 | ICLR | Building Math Agents with Multi-Turn Iterative Preference Learning. | Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg, Zhen Qin, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, Chi Jin, Tong Zhang, Tianqi Liu |
| 2025 | ICLR | Learning from negative feedback, or positive feedback or both. | Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari, Jost Tobias Springenberg, Tim Hertweck, Michael Bloesch, Rishabh Joshi, Thomas Lampe, Junhyuk Oh, Nicolas Heess, Jonas Buchli, Martin A. Riedmiller |
| 2024 | AISTATS | A General Theoretical Paradigm to Understand Learning from Human Preferences. | Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello |
| 2024 | ICLR | Unlocking the Power of Representations in Long-term Novelty-based Exploration. | Alaa Saade, Steven Kapturowski, Daniele Calandriello, Charles Blundell, Pablo Sprechmann, Leopoldo Sarra, Oliver Groth, Michal Valko, Bilal Piot |
| 2024 | ICML | Human Alignment of Large Language Models through Online Preference Optimisation. | Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot |
| 2024 | ICML | Nash Learning from Human Feedback. | Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot |
| 2024 | ICML | Generalized Preference Optimization: A Unified Approach to Offline Alignment. | Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot |
| 2023 | ICML | The Edge of Orthogonality: A Simple View of What Makes BYOL Tick. | Pierre Harvey Richemond, Allison C. Tam, Yunhao Tang, Florian Strub, Bilal Piot, Felix Hill |
| 2023 | ICML | Understanding Self-Predictive Learning for Reinforcement Learning. | Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko |
| 2022 | ICLR | Emergent Communication at Scale. | Rahma Chaabouni, Florian Strub, Florent Altch, Eugene Tarassov, Corentin Tallec, Elnaz Davoodi, Kory Wallace Mathewson, Olivier Tieleman, Angeliki Lazaridou, Bilal Piot |
| 2020 | ICLR | Never Give Up: Learning Directed Exploration Strategies. | Adri Puigdomnech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martn Arjovsky, Alexander Pritzel, Andrew Bolt, Charles Blundell |
| 2020 | ICML | Agent57: Outperforming the Atari Human Benchmark. | Adri Puigdomnech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Charles Blundell |
| 2020 | ICML | Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning. | Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar |
| 2018 | AAAI | Rainbow: Combining Improvements in Deep Reinforcement Learning. | Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, David Silver |
| 2018 | AAAI | Deep Q-learning From Demonstrations. | Todd Hester, Matej Vecerk, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, Audrunas Gruslys |
| 2018 | AISTATS | Actor-Critic Fictitious Play in Simultaneous Move Multistage Games. | Julien Prolat, Bilal Piot, Olivier Pietquin |
| 2018 | ICLR | Noisy Networks For Exploration. | Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg |
| 2018 | ICLR | The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning. | Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos |
| 2017 | AISTATS | Learning Nash Equilibrium for General-Sum Markov Games from Batch Data. | Julien Prolat, Florian Strub, Bilal Piot, Olivier Pietquin |
| 2017 | IJCAI | End-to-end optimization of goal-driven and visually grounded dialogue systems. | Florian Strub, Harm de Vries, Jrmie Mary, Bilal Piot, Aaron C. Courville, Olivier Pietquin |
| 2016 | AISTATS | On the Use of Non-Stationary Strategies for Solving Two-Player Zero-Sum Markov Games. | Julien Prolat, Bilal Piot, Bruno Scherrer, Olivier Pietquin |
| 2016 | ICML | Softened Approximate Policy Iteration for Markov Games. | Julien Prolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, Olivier Pietquin |
| 2015 | ICML | Approximate Dynamic Programming for Two-Player Zero-Sum Markov Games. | Julien Prolat, Bruno Scherrer, Bilal Piot, Olivier Pietquin |
| 2015 | ICML | Imitation Learning Applied to Embodied Conversational Agents. | Bilal Piot, Olivier Pietquin, Matthieu Geist |
| 2015 | IJCAI | Inverse Reinforcement Learning in Relational Domains. | Thibaut Munzer, Bilal Piot, Matthieu Geist, Olivier Pietquin, Manuel Lopes |
| 2014 | Interspeech | Predicting when to laugh with structured classification. | Bilal Piot, Olivier Pietquin, Matthieu Geist |