Central Path Proximal Policy Optimization
Published · Exploration in AI Today Workshop at ICML 2025. In constrained Markov decision processes, enforcing constraints during training is often thought of as decreasing the final return. Recently, it was shown that constraints can be incorporated directly into the policy geometry,…
