Skip to content

Odalric-Ambrym Maillard

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

21

Venues

6

Active years

2009–2025

Best venue rank

A*

Where they publish

Papers

21 indexed papers, newest first.

YearVenueTitleAuthors
2025ICMLMonte-Carlo Tree Search with Uncertainty Propagation via Optimal Transport.Tuan Dam, Pascal Stenger, Lukas Schneider, Joni Pajarinen, Carlo D'Eramo, Odalric-Ambrym Maillard
2024ALTCRIMED: Lower and Upper Bounds on Regret for Bandits with Unbounded Stochastic Corruption.Shubhada Agrawal, Timothe Mathieu, Debabrota Basu, Odalric-Ambrym Maillard
2024UAIPower Mean Estimation in Stochastic Monte-Carlo Tree Search.Tuan Dam, Odalric-Ambrym Maillard, Emilie Kaufmann
2023ACMLLogarithmic regret in communicating MDPs: Leveraging known dynamics with bandits.Hassan Saber, Fabien Pesquerel, Odalric-Ambrym Maillard, Mohammad Sadegh Talebi
2023AISTATSExploration in Reward Machines with Low Regret.Hippolyte Bourel, Anders Jonsson, Odalric-Ambrym Maillard, Mohammad Sadegh Talebi
2021AISTATSReinforcement Learning in Parametric MDPs with Exponential Families.Sayak Ray Chowdhury, Aditya Gopalan, Odalric-Ambrym Maillard
2021ICLRLearning Value Functions in Deep Policy Gradients using Residual Variance.Yannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe Preux
2020ACMLMonte-Carlo Graph Search: the Value of Merging Similar States.Edouard Leurent, Odalric-Ambrym Maillard
2019ACMLModel-Based Reinforcement Learning Exploiting State-Action Equivalence.Mahsa Asadi, Mohammad Sadegh Talebi, Hippolyte Bourel, Odalric-Ambrym Maillard
2019ALTSequential change-point detection: Laplace concentration of scan statistics and non-asymptotic delay bounds.Odalric-Ambrym Maillard
2018ALTVariance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs.Mohammad Sadegh Talebi, Odalric-Ambrym Maillard
2017ALTBoundary Crossing for General Exponential Families.Odalric-Ambrym Maillard
2017ALTEfficient tracking of a growing number of experts.Jaouad Mourtada, Odalric-Ambrym Maillard
2017ICMLSpectral Learning from a Single Trajectory under Finite-State Policies.Borja Balle, Odalric-Ambrym Maillard
2016ICMLPliable Rejection Sampling.Akram Erraqabi, Michal Valko, Alexandra Carpentier, Odalric-Ambrym Maillard
2014ALTSelecting Near-Optimal Approximate State Representations in Reinforcement Learning.Ronald Ortner, Odalric-Ambrym Maillard, Daniil Ryabko
2014ICMLLatent Bandits.Odalric-Ambrym Maillard, Shie Mannor
2013AISTATSCompeting with an Infinite Set of Models in Reinforcement Learning.Phuong Nguyen, Odalric-Ambrym Maillard, Daniil Ryabko, Ronald Ortner
2013ALTRobust Risk-Averse Stochastic Multi-armed Bandits.Odalric-Ambrym Maillard
2013ICMLOptimal Regret Bounds for Selecting the State Representation in Reinforcement Learning.Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, Daniil Ryabko
2009ALTComplexity versus Agreement for Many Views.Odalric-Ambrym Maillard, Nicolas Vayatis