| 2026 | COLT | Sharp analysis of linear ensemble sampling. | David Janz, Arya Akhavan, Csaba Szepesvri |
| 2026 | COLT | Continuous time policy evaluation is easier with noisy dynamics. | Samuel Robertson, Thomas Newton, Csaba Szepesvri |
| 2026 | COLT | Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear q | Volodymyr Tkachuk, Csaba Szepesvri, Xiaoqi Tan |
| 2025 | COLT | Thompson Sampling for Bandit Convex Optimisation. | Alireza Bakhtiari, Tor Lattimore, Csaba Szepesvri |
| 2024 | AISTATS | Exploration via linearly perturbed loss minimisation. | David Janz, Shuai Liu, Alex Ayoub, Csaba Szepesvri |
| 2024 | ICLR | Stochastic Gradient Descent for Gaussian Processes Done Right. | Jihao Andreas Lin, Shreyas Padhy, Javier Antorn, Austin Tripp, Alexander Terenin, Csaba Szepesvri, Jos Miguel Hernndez-Lobato, David Janz |
| 2024 | ICML | Switching the Loss Reduces the Cost in Batch Reinforcement Learning. | Alex Ayoub, Kaiwen Wang, Vincent Liu, Samuel Robertson, James McInerney, Dawen Liang, Nathan Kallus, Csaba Szepesvri |
| 2023 | AISTATS | Efficient Planning in Combinatorial Action Spaces with Applications to Cooperative Multi-Agent Reinforcement Learning. | Volodymyr Tkachuk, Seyed Alireza Bakhtiari, Johannes Kirschner, Matej Jusup, Ilija Bogunovic, Csaba Szepesvri |
| 2023 | COLT | Exponential Hardness of Reinforcement Learning with Linear Function Approximation. | Sihan Liu, Gaurav Mahajan, Daniel Kane, Shachar Lovett, Gellrt Weisz, Csaba Szepesvri |
| 2023 | ICLR | Optimistic Exploration with Learned Features Provably Solves Markov Decision Processes with Neural Dynamics. | Sirui Zheng, Lingxiao Wang, Shuang Qiu, Zuyue Fu, Zhuoran Yang, Csaba Szepesvri, Zhaoran Wang |
| 2023 | ICML | The Optimal Approximation Factors in Misspecified Off-Policy Value Function Estimation. | Philip Amortila, Nan Jiang, Csaba Szepesvri |
| 2023 | ICML | Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice. | Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo |
| 2023 | ICML | Stochastic Gradient Succeeds for Bandits. | Jincheng Mei, Zixin Zhong, Bo Dai, Alekh Agarwal, Csaba Szepesvri, Dale Schuurmans |
| 2023 | ICML | Revisiting Simple Regret: Fast Rates for Returning a Good Arm. | Yao Zhao, Connor Stephens, Csaba Szepesvri, Kwang-Sung Jun |
| 2023 | STOC | Optimistic MLE: A Generic Model-Based Algorithm for Partially Observable Sequential Decision Making. | Qinghua Liu, Praneeth Netrapalli, Csaba Szepesvri, Chi Jin |
| 2022 | AISTATS | Confident Least Square Value Iteration with Local Access to a Simulator. | Botao Hao, Nevena Lazic, Dong Yin, Yasin Abbasi-Yadkori, Csaba Szepesvri |
| 2022 | AISTATS | Faster Rates, Adaptive Algorithms, and Finite-Time Bounds for Linear Composition Optimization and Gradient TD Learning. | Anant Raj, Pooria Joulani, Andrs Gyrgy, Csaba Szepesvri |
| 2022 | AISTATS | The Curse of Passive Data Collection in Batch Reinforcement Learning. | Chenjun Xiao, Ilbin Lee, Bo Dai, Dale Schuurmans, Csaba Szepesvri |
| 2022 | ALT | TensorPlan and the Few Actions Lower Bound for Planning in MDPs under Linear Realizability of Optimal Value Functions. | Gellrt Weisz, Csaba Szepesvri, Andrs Gyrgy |
| 2022 | ALT | Efficient local planning with linear function approximation. | Dong Yin, Botao Hao, Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvri |
| 2022 | COLT | When Is Partially Observable Reinforcement Learning Not Scary? | Qinghua Liu, Alan Chung, Csaba Szepesvri, Chi Jin |
| 2022 | UAI | A free lunch from the noise: Provable and practical exploration for representation learning. | Tongzheng Ren, Tianjun Zhang, Csaba Szepesvri, Bo Dai |
| 2021 | AISTATS | Adaptive Approximate Policy Iteration. | Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori, Pooria Joulani, Csaba Szepesvri |
| 2021 | AISTATS | Online Sparse Reinforcement Learning. | Botao Hao, Tor Lattimore, Csaba Szepesvri, Mengdi Wang |
| 2021 | AISTATS | Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting. | Ilja Kuzborskij, Claire Vernade, Andrs Gyrgy, Csaba Szepesvri |
| 2021 | ALT | Exponential Lower Bounds for Planning in MDPs With Linearly-Realizable Optimal Action-Value Functions. | Gellrt Weisz, Philip Amortila, Csaba Szepesvri |
| 2021 | COLT | Asymptotically Optimal Information-Directed Sampling. | Johannes Kirschner, Tor Lattimore, Claire Vernade, Csaba Szepesvri |
| 2021 | COLT | Nonparametric Regression with Shallow Overparameterized Neural Networks Trained by GD with Early Stopping. | Ilja Kuzborskij, Csaba Szepesvri |
| 2021 | COLT | On Query-efficient Planning in MDPs under Linear Realizability of the Optimal State-value Function. | Gellrt Weisz, Philip Amortila, Barnabs Janzer, Yasin Abbasi-Yadkori, Nan Jiang, Csaba Szepesvri |
| 2021 | COLT | Nearly Minimax Optimal Reinforcement Learning for Linear Mixture Markov Decision Processes. | Dongruo Zhou, Quanquan Gu, Csaba Szepesvri |
| 2021 | ICML | Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient. | Botao Hao, Yaqi Duan, Tor Lattimore, Csaba Szepesvri, Mengdi Wang |
| 2021 | ICML | Bootstrapping Fitted Q-Evaluation for Off-Policy Inference. | Botao Hao, Xiang Ji, Yaqi Duan, Hao Lu, Csaba Szepesvri, Mengdi Wang |
| 2021 | ICML | A Distribution-dependent Analysis of Meta Learning. | Mikhail Konobeev, Ilja Kuzborskij, Csaba Szepesvri |
| 2021 | ICML | Meta-Thompson Sampling. | Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-Wei Hsu, Martin Mladenov, Craig Boutilier, Csaba Szepesvri |
| 2021 | ICML | Improved Regret Bound and Experience Replay in Regularized Policy Iteration. | Nevena Lazic, Dong Yin, Yasin Abbasi-Yadkori, Csaba Szepesvri |
| 2021 | ICML | Leveraging Non-uniformity in First-order Non-convex Optimization. | Jincheng Mei, Yue Gao, Bo Dai, Csaba Szepesvri, Dale Schuurmans |
| 2021 | ICML | On the Optimality of Batch Policy Optimization Algorithms. | Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvri, Dale Schuurmans |
| 2020 | AISTATS | Adaptive Exploration in Linear Contextual Bandit. | Botao Hao, Tor Lattimore, Csaba Szepesvri |
| 2020 | AISTATS | Randomized Exploration in Generalized Linear Bandits. | Branislav Kveton, Manzil Zaheer, Csaba Szepesvri, Lihong Li, Mohammad Ghavamzadeh, Craig Boutilier |
| 2020 | COLT | Exploration by Optimisation in Partial Monitoring. | Tor Lattimore, Csaba Szepesvri |
| 2020 | ICLR | Behaviour Suite for Reinforcement Learning. | Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvri, Satinder Singh, Benjamin Van Roy, Richard S. Sutton, David Silver, Hado van Hasselt |
| 2020 | ICML | Model-Based Reinforcement Learning with Value-Targeted Regression. | Alex Ayoub, Zeyu Jia, Csaba Szepesvri, Mengdi Wang, Lin Yang |
| 2020 | ICML | A simpler approach to accelerated optimization: iterative averaging meets optimism. | Pooria Joulani, Anant Raj, Andrs Gyrgy, Csaba Szepesvri |
| 2020 | ICML | Learning with Good Feature Representations in Bandits and in RL with a Generative Model. | Tor Lattimore, Csaba Szepesvri, Gellrt Weisz |
| 2020 | ICML | On the Global Convergence Rates of Softmax Policy Gradient Methods. | Jincheng Mei, Chenjun Xiao, Csaba Szepesvri, Dale Schuurmans |
| 2019 | AAAI | An Exponential Tail Bound for the Deleted Estimate. | Karim T. Abou-Moustafa, Csaba Szepesvri |
| 2019 | AISTATS | Model-Free Linear Quadratic Control via Reduction to Expert Prediction. | Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvri |
| 2019 | AISTATS | Online Algorithm for Unsupervised Sensor Selection. | Arun Verma, Manjesh Kumar Hanawal, Csaba Szepesvri, Venkatesh Saligrama |
| 2019 | ALT | An Exponential Efron-Stein Inequality for | Karim T. Abou-Moustafa, Csaba Szepesvri |
| 2019 | ALT | Cleaning up the neighborhood: A full classification for adversarial partial monitoring. | Tor Lattimore, Csaba Szepesvri |
| 2019 | COLT | Distribution-Dependent Analysis of Gibbs-ERM Principle. | Ilja Kuzborskij, Nicol Cesa-Bianchi, Csaba Szepesvri |
| 2019 | COLT | An Information-Theoretic Approach to Minimax Regret in Partial Monitoring. | Tor Lattimore, Csaba Szepesvri |
| 2019 | ICLR | Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures. | Jonathan Uesato, Ananya Kumar, Csaba Szepesvri, Tom Erez, Avraham Ruderman, Keith Anderson, Krishnamurthy (Dj) Dvijotham, Nicolas Heess, Pushmeet Kohli |
| 2019 | ICML | Garbage In, Reward Out: Bootstrapping Exploration in Multi-Armed Bandits. | Branislav Kveton, Csaba Szepesvri, Sharan Vaswani, Zheng Wen, Tor Lattimore, Mohammad Ghavamzadeh |
| 2019 | ICML | Online Learning to Rank with Features. | Shuai Li, Tor Lattimore, Csaba Szepesvri |
| 2019 | ICML | CapsAndRuns: An Improved Method for Approximately Optimal Algorithm Configuration. | Gellrt Weisz, Andrs Gyrgy, Csaba Szepesvri |
| 2019 | IJCAI | Perturbed-History Exploration in Stochastic Multi-Armed Bandits. | Branislav Kveton, Csaba Szepesvri, Mohammad Ghavamzadeh, Craig Boutilier |
| 2019 | UAI | BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback. | Chang Li, Branislav Kveton, Tor Lattimore, Ilya Markov, Maarten de Rijke, Csaba Szepesvri, Masrour Zoghi |
| 2019 | UAI | Perturbed-History Exploration in Stochastic Linear Bandits. | Branislav Kveton, Csaba Szepesvri, Mohammad Ghavamzadeh, Craig Boutilier |
| 2018 | AISTATS | Linear Stochastic Approximation: How Far Does Constant Step-Size and Iterate Averaging Go? | Chandrashekar Lakshminarayanan, Csaba Szepesvri |
| 2018 | ICML | Gradient Descent for Sparse Rank-One Matrix Completion for Crowd-Sourced Aggregation of Sparsely Interacting Workers. | Yao Ma, Alexander Olshevsky, Csaba Szepesvri, Venkatesh Saligrama |
| 2018 | ICML | Bandits with Delayed, Aggregated Anonymous Feedback. | Ciara Pike-Burke, Shipra Agrawal, Csaba Szepesvri, Steffen Grnewlder |
| 2018 | ICML | LEAPSANDBOUNDS: A Method for Approximately Optimal Algorithm Configuration. | Gellrt Weisz, Andrs Gyrgy, Csaba Szepesvri |
| 2018 | ISAIM | An Exponential Tail Bound for Lq Stable Learning Rules. Application to k-Folds Cross-Validation. | Karim T. Abou-Moustafa, Csaba Szepesvri |
| 2017 | AISTATS | Unsupervised Sequential Sensor Acquisition. | Manjesh Kumar Hanawal, Csaba Szepesvri, Venkatesh Saligrama |
| 2017 | AISTATS | Stochastic Rank-1 Bandits. | Sumeet Katariya, Branislav Kveton, Csaba Szepesvri, Claire Vernade, Zheng Wen |
| 2017 | AISTATS | The End of Optimism? An Asymptotic Analysis of Finite-Armed Linear Bandits. | Tor Lattimore, Csaba Szepesvri |
| 2017 | ALT | Structured Best Arm Identification with Fixed Confidence. | Ruitong Huang, Mohammad M. Ajallooeian, Csaba Szepesvri, Martin Mller |
| 2017 | ALT | A Modular Analysis of Adaptive (Non-)Convex Optimization: Optimism, Composite Objectives, and Variational Bounds. | Pooria Joulani, Andrs Gyrgy, Csaba Szepesvri |
| 2017 | ICML | Online Learning to Rank in Stochastic Click Models. | Masrour Zoghi, Toms Tunys, Mohammad Ghavamzadeh, Branislav Kveton, Csaba Szepesvri, Zheng Wen |
| 2017 | IJCAI | Bernoulli Rank-1 Bandits for Click Feedback. | Sumeet Katariya, Branislav Kveton, Csaba Szepesvri, Claire Vernade, Zheng Wen |
| 2016 | AAAI | Delay-Tolerant Online Convex Optimization: Unified Analysis and Adaptive-Gradient Algorithms. | Pooria Joulani, Andrs Gyrgy, Csaba Szepesvri |
| 2016 | AAAI | Compressed Conditional Mean Embeddings for Model-Based Reinforcement Learning. | Guy Lever, John Shawe-Taylor, Ronnie Stafford, Csaba Szepesvri |
| 2016 | AISTATS | (Bandit) Convex Optimization with Biased Noisy Gradient Oracles. | Xiaowei Hu, Prashanth L. A., Andrs Gyrgy, Csaba Szepesvri |
| 2016 | ICML | Cumulative Prospect Theory Meets Reinforcement Learning: Prediction and Control. | Prashanth L. A., Cheng Jie, Michael C. Fu, Steven I. Marcus, Csaba Szepesvri |
| 2016 | ICML | Shifting Regret, Mirror Descent, and Matrices. | Andrs Gyrgy, Csaba Szepesvri |
| 2016 | ICML | DCM Bandits: Learning to Rank with Multiple Clicks. | Sumeet Katariya, Branislav Kveton, Csaba Szepesvri, Zheng Wen |
| 2016 | ICML | Conservative Bandits. | Yifan Wu, Roshan Shariff, Tor Lattimore, Csaba Szepesvri |
| 2015 | AAAI | Decision-Theoretic Clustering of Strategies. | Nolan Bard, Deon Nicholas, Csaba Szepesvri, Michael Bowling |
| 2015 | AAAI | Pathological Effects of Variance on Classification-Based Policy Iteration. | Bernardo vila Pires, Csaba Szepesvri |
| 2015 | AISTATS | Near-optimal max-affine estimators for convex regression. | Gbor Balzs, Andrs Gyrgy, Csaba Szepesvri |
| 2015 | AISTATS | Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits. | Branislav Kveton, Zheng Wen, Azin Ashkan, Csaba Szepesvri |
| 2015 | AISTATS | Toward Minimax Off-policy Value Estimation. | Lihong Li, Rmi Munos, Csaba Szepesvri |
| 2015 | AISTATS | Exploiting Symmetries to Construct Efficient MCMC Algorithms With an Application to SLAM. | Roshan Shariff, Andrs Gyrgy, Csaba Szepesvri |
| 2015 | ICML | Deterministic Independent Component Analysis. | Ruitong Huang, Andrs Gyrgy, Csaba Szepesvri |
| 2015 | ICML | Cascading Bandits: Learning to Rank in the Cascade Model. | Branislav Kveton, Csaba Szepesvri, Zheng Wen, Azin Ashkan |
| 2015 | ICML | On Identifying Good Options under Combinatorially Structured Feedback in Finite Noisy Environments. | Yifan Wu, Andrs Gyrgy, Csaba Szepesvri |
| 2015 | IJCAI | Fast Cross-Validation for Incremental Learning. | Pooria Joulani, Andrs Gyrgy, Csaba Szepesvri |
| 2015 | UAI | Bayesian Optimal Control of Smoothly Parameterized Systems. | Yasin Abbasi-Yadkori, Csaba Szepesvri |
| 2014 | AISTATS | A Finite-Sample Generalization Bound for Semiparametric Regression: Partially Linear Models. | Ruitong Huang, Csaba Szepesvri |
| 2014 | ALT | On Learning the Optimal Waiting Time. | Tor Lattimore, Andrs Gyrgy, Csaba Szepesvri |
| 2014 | ICML | Online Learning in Markov Decision Processes with Changing Cost Sequences. | Travis Dick, Andrs Gyrgy, Csaba Szepesvri |
| 2014 | ICML | Adaptive Monte Carlo via Bandit Allocation. | James Neufeld, Andrs Gyrgy, Csaba Szepesvri, Dale Schuurmans |
| 2014 | ISAIM | Generalization Bounds for Partially Linear Models. | Ruitong Huang, Csaba Szepesvri |
| 2014 | UAI | Optimal Resource Allocation with Semi-Bandit Feedback. | Tor Lattimore, Koby Crammer, Csaba Szepesvri |
| 2013 | ICML | A Randomized Mirror Descent Algorithm for Large Scale Multiple Kernel Learning. | Arash Afkanpour, Andrs Gyrgy, Csaba Szepesvri, Michael Bowling |
| 2013 | ICML | Online Learning under Delayed Feedback. | Pooria Joulani, Andrs Gyrgy, Csaba Szepesvri |
| 2013 | ICML | Cost-sensitive Multiclass Classification Risk Bounds. | Bernardo vila Pires, Csaba Szepesvri, Mohammad Ghavamzadeh |
| 2013 | ICML | Characterizing the Representer Theorem. | Yaoliang Yu, Hao Cheng, Dale Schuurmans, Csaba Szepesvri |
| 2012 | AAAI | Approximate Policy Iteration with Linear Action Models. | Hengshuai Yao, Csaba Szepesvri |
| 2012 | ALT | Partial Monitoring with Side Information. | Gbor Bartk, Csaba Szepesvri |
| 2012 | ICML | An adaptive algorithm for finite stochastic partial monitoring. | Gbor Bartk, Navid Zolghadr, Csaba Szepesvri |
| 2012 | ICML | Statistical linear estimation with penalized estimators: an application to reinforcement learning. | Bernardo vila Pires, Csaba Szepesvri |
| 2012 | ICML | Analysis of Kernel Mean Matching under Covariate Shift. | Yaoliang Yu, Csaba Szepesvri |
| 2011 | ALT | Editors' Introduction. | Jyrki Kivinen, Csaba Szepesvri, Esko Ukkonen, Thomas Zeugmann |
| 2011 | INFOCOM | Sequential learning for optimal monitoring of multi-channel wireless networks. | Pallavi Arora, Csaba Szepesvri, Rong Zheng |
| 2011 | UAI | PAC-Bayesian Policy Evaluation for Reinforcement Learning. | Mahdi Milani Fard, Joelle Pineau, Csaba Szepesvri |
| 2010 | ALT | Toward a Classification of Finite Partial-Monitoring Games. | Gbor Bartk, Dvid Pl, Csaba Szepesvri |
| 2010 | COLT | The Online Loop-free Stochastic Shortest-Path Problem. | Gergely Neu, Andrs Gyrgy, Csaba Szepesvri |
| 2010 | ICML | Budgeted Distribution Learning of Belief Net Parameters. | Liuyang Li, Barnabs Pczos, Csaba Szepesvri, Russell Greiner |
| 2010 | ICML | Toward Off-Policy Learning Control with Function Approximation. | Hamid Reza Maei, Csaba Szepesvri, Shalabh Bhatnagar, Richard S. Sutton |
| 2010 | ICML | Model-based reinforcement learning with nearly tight exploration complexity bounds. | Istvan Szita, Csaba Szepesvri |
| 2010 | IROS | Extending rapidly-exploring random trees for asymptotically optimal anytime motion planning. | Yasin Abbasi-Yadkori, Joseph Modayil, Csaba Szepesvri |
| 2009 | ICML | Workshop summary: On-line learning with limited feedback. | Jean-Yves Audibert, Peter Auer, Alessandro Lazaric, Rmi Munos, Daniil Ryabko, Csaba Szepesvri |
| 2009 | ICML | Learning to segment from a few well-selected training images. | Alireza Farhangfar, Russell Greiner, Csaba Szepesvri |
| 2009 | ICML | Learning when to stop thinking and do something! | Barnabs Pczos, Yasin Abbasi-Yadkori, Csaba Szepesvri, Russell Greiner, Nathan R. Sturtevant |
| 2009 | ICML | Fast gradient-descent methods for temporal-difference learning with linear function approximation. | Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvri, Eric Wiewiora |
| 2009 | ICRA | Model-based and model-free reinforcement learning for visual servoing. | Amir Massoud Farahmand, Azad Shademan, Martin Jgersand, Csaba Szepesvri |
| 2008 | ALT | Active Learning in Multi-armed Bandits. | Andrs Antos, Varun Grover, Csaba Szepesvri |
| 2008 | ALT | Active Learning of Group-Structured Environments. | Gbor Bartk, Csaba Szepesvri, Sandra Zilles |
| 2008 | ICML | Empirical Bernstein stopping. | Volodymyr Mnih, Csaba Szepesvri, Jean-Yves Audibert |
| 2008 | UAI | Speeding Up Planning in Markov Decision Processes via Automatically Constructed Abstraction. | Alejandro Isaza, Csaba Szepesvri, Vadim Bulitko, Russell Greiner |
| 2008 | UAI | Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping. | Richard S. Sutton, Csaba Szepesvri, Alborz Geramifard, Michael H. Bowling |
| 2007 | ALT | Tuning Bandit Algorithms in Stochastic Environments. | Jean-Yves Audibert, Rmi Munos, Csaba Szepesvri |
| 2007 | COLT | Improved Rates for the Stochastic Continuum-Armed Bandit Problem. | Peter Auer, Ronald Ortner, Csaba Szepesvri |
| 2007 | ICML | Manifold-adaptive dimension estimation. | Amir Massoud Farahmand, Csaba Szepesvri, Jean-Yves Audibert |
| 2007 | IJCAI | Sequence Prediction Exploiting Similary Information. | Istvn Br, Zoltn Szamonek, Csaba Szepesvri |
| 2007 | IJCAI | Continuous Time Associative Bandit Problems. | Andrs Gyrgy, Levente Kocsis, Ivett Szab, Csaba Szepesvri |
| 2007 | UAI | Apprenticeship Learning using Inverse Reinforcement Learning and Gradient Methods. | Gergely Neu, Csaba Szepesvri |
| 2006 | COLT | Learning Near-Optimal Policies with Bellman-Residual Minimization Based Fitted Policy Iteration and a Single Sample Path. | Andrs Antos, Csaba Szepesvri, Rmi Munos |
| 2005 | ICDM | X-mHMM: An Efficient Algorithm for Training Mixtures of HMMs When the Number of Mixtures Is Unknown. | Zoltn Szamonek, Csaba Szepesvri |
| 2005 | ICML | Finite time bounds for sampling based fitted value iteration. | Csaba Szepesvri, Rmi Munos |
| 2004 | AAAI | Shortest Path Discovery Problems: A Framework, Algorithms and Experimental Results. | Csaba Szepesvri |
| 2004 | ECAI | Kernel Machine Based Feature Extraction Algorithms for Regression Problems. | Csaba Szepesvri, Andrs Kocsor, Kornl Kovcs |
| 2004 | ECCV | Enhancing Particle Filters Using Local Likelihood Sampling. | Pter Torma, Csaba Szepesvri |
| 2004 | ICML | Interpolation-based Q-learning. | Csaba Szepesvri, William D. Smart |
| 2003 | AISTATS | Sequential Importance Sampling for Visual Tracking Reconsidered. | Pter Torma, Csaba Szepesvri |
| 1998 | ICML | Multi-criteria Reinforcement Learning. | Zoltn Gbor, Zsolt Kalmr, Csaba Szepesvri |
| 1996 | ICANN | Inverse Dynamics Controllers for Robust Control: Consequences for Neurocontrollers. | Csaba Szepesvri, Andrs Lrincz |
| 1996 | ICML | A Generalized Reinforcement-Learning Model: Convergence and Applications. | Michael L. Littman, Csaba Szepesvri |