Skip to content

Policy Optimization for CMDPs with Bandit Feedback: Learning Stochastic and Adversarial Constraints.

Francesco Emanuele Stradi, Anna Lunghi, Matteo Castiglioni, Alberto Marchesi, Nicola Gatti

VenueA*ICML
Year2025
ProceedingsICML

Browse the full ICML paper archive.