| 2025 | AAAI | Data Augmentation for Instruction Following Policies via Trajectory Segmentation. | Niklas Hpner, Ilaria Tiddi, Herke van Hoof |
| 2024 | CIKM | Mitigating Exposure Bias in Online Learning to Rank Recommendation: A Novel Reward Model for Cascading Bandits. | Masoud Mansoury, Bamshad Mobasher, Herke van Hoof |
| 2024 | ICAPS | Planning with a Learned Policy Basis to Optimally Solve Complex Tasks. | David Kuric, Guillermo Infante, Vicen Gmez, Anders Jonsson, Herke van Hoof |
| 2024 | SIGIR | Going Beyond Popularity and Positivity Bias: Correcting for Multifactorial Bias in Recommender Systems. | Jin Huang, Harrie Oosterhuis, Masoud Mansoury, Herke van Hoof, Maarten de Rijke |
| 2023 | ICLR | Bridge the Inference Gaps of Neural Processes via Expectation Maximization. | Qi Wang, Marco Federici, Herke van Hoof |
| 2022 | AAAI | Fast and Data Efficient Reinforcement Learning from Pixels via Non-parametric Value Approximation. | Alexander Long, Alan Blair, Herke van Hoof |
| 2022 | CPAIOR | Deep Policy Dynamic Programming for Vehicle Routing Problems. | Wouter Kool, Herke van Hoof, Joaquim A. S. Gromicho, Max Welling |
| 2022 | ICLR | Multi-Agent MDP Homomorphic Networks. | Elise van der Pol, Herke van Hoof, Frans A. Oliehoek, Max Welling |
| 2022 | ICML | Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models and Amortized Policy Search. | Qi Wang, Herke van Hoof |
| 2022 | IJCAI | Leveraging Class Abstraction for Commonsense Reinforcement Learning via Residual Policy Gradient Methods. | Niklas Hpner, Ilaria Tiddi, Herke van Hoof |
| 2022 | IJCAI | Value Refinement Network (VRN). | Jan Whlke, Felix Schmitt, Herke van Hoof |
| 2022 | IJCNN | Logic-based AI for Interpretable Board Game Winner Prediction with Tsetlin Machine. | Charul Giri, Ole-Christoffer Granmo, Herke van Hoof, Christian D. Blakely |
| 2021 | ICML | Deep Coherent Exploration for Continuous Control. | Yijie Zhang, Herke van Hoof |
| 2021 | ICRA | Hierarchies of Planning and Reinforcement Learning for Robot Navigation. | Jan Whlke, Felix Schmitt, Herke van Hoof |
| 2020 | ICANN | Social Navigation with Human Empowerment Driven Deep Reinforcement Learning. | Tessa van der Heiden, Florian Mirus, Herke van Hoof |
| 2020 | ICLR | Estimating Gradients for Discrete Random Variables by Sampling without Replacement. | Wouter Kool, Herke van Hoof, Max Welling |
| 2020 | ICML | Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent Variables. | Qi Wang, Herke van Hoof |
| 2020 | RecSys | Keeping Dataset Biases out of the Simulation: A Debiased Simulator for Reinforcement Learning based Recommender Systems. | Jin Huang, Harrie Oosterhuis, Maarten de Rijke, Herke van Hoof |
| 2019 | ICLR | Attention, Learn to Solve Routing Problems! | Wouter Kool, Herke van Hoof, Max Welling |
| 2019 | ICLR | Buy 4 REINFORCE Samples, Get a Baseline for Free! | Wouter Kool, Herke van Hoof, Max Welling |
| 2019 | ICML | Stochastic Beams and Where To Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement. | Wouter Kool, Herke van Hoof, Max Welling |
| 2019 | IROS | Deep Generative Modeling of LiDAR Data. | Lucas Caccia, Herke van Hoof, Aaron C. Courville, Joelle Pineau |
| 2019 | ICRA | Uncertainty Aware Learning from Demonstrations in Multiple Contexts using Bayesian Neural Networks. | Sanjay Thakur, Herke van Hoof, Juan Camilo Gamboa Higuera, Doina Precup, David Meger |
| 2018 | EMNLP | BanditSum: Extractive Summarization as a Contextual Bandit. | Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, Jackie Chi Kit Cheung |
| 2018 | ICML | Addressing Function Approximation Error in Actor-Critic Methods. | Scott Fujimoto, Herke van Hoof, David Meger |
| 2018 | ICML | An Inference-Based Policy Gradient Method for Learning Options. | Matthew J. A. Smith, Herke van Hoof, Joelle Pineau |
| 2018 | ICRA | Eager and Memory-Based Non-Parametric Stochastic Search Methods for Learning Control. | Victor Barbaros, Herke van Hoof, Abbas Abdolmaleki, David Meger |
| 2017 | AAAI | Policy Search with High-Dimensional Context Variables. | Voot Tangkaratt, Herke van Hoof, Simone Parisi, Gerhard Neumann, Jan Peters, Masashi Sugiyama |
| 2016 | IROS | Stable reinforcement learning with autoencoders for tactile and visual data. | Herke van Hoof, Nutan Chen, Maximilian Karl, Patrick van der Smagt, Jan Peters |
| 2016 | IROS | Active tactile object exploration with Gaussian processes. | Zhengkun Yi, Roberto Calandra, Filipe Veiga, Herke van Hoof, Tucker Hermans, Yilei Zhang, Jan Peters |
| 2015 | AISTATS | Learning of Non-Parametric Control Policies with High-Dimensional State Features. | Herke van Hoof, Jan Peters, Gerhard Neumann |
| 2015 | ICRA | Towards learning hierarchical skills for multi-phase manipulation tasks. | Oliver Kroemer, Christian Daniel, Gerhard Neumann, Herke van Hoof, Jan Peters |
| 2015 | IROS | Stabilizing novel objects by learning to predict tactile slip. | Filipe Veiga, Herke van Hoof, Jan Peters, Tucker Hermans |
| 2014 | ICRA | Policy search for learning robot control using sparse data. | Bastian Bischoff, Duy Nguyen-Tuong, Herke van Hoof, Andrew McHutchon, Carl E. Rasmussen, Alois C. Knoll, Jan Peters, Marc Peter Deisenroth |
| 2014 | ICRA | Learning to predict phases of manipulation tasks as hidden states. | Oliver Kroemer, Herke van Hoof, Gerhard Neumann, Jan Peters |
| 2012 | IROS | Maximally informative interaction learning for scene exploration. | Herke van Hoof, Oliver Kroemer, Heni Ben Amor, Jan Peters |