Robust Peak-cost Constrained Reinforcement Learning
Researchers introduce a robust peak-cost constrained reinforcement learning (RL) framework that aims to control the maximum cost encountered along a trajectory, addressing the inability of standard constrained Markov decision processes (CMDPs) to handle catastrophic single violations. They demonstrate that peak-cost MDPs may not have zero duality gap and propose a surrogate optimization method with robust value estimation to address this. Experimental results indicate that their approach enforces safety constraints effectively even under dynamics perturbations.
Why it matters: This work advances safe RL by providing a method to handle catastrophic single violations, which is crucial for safety-critical applications.
Full story at: arXiv Machine Learning ↗