Constrained Policy Optimization
In reinforcement learning (RL), agents learn to act by trial and error, gradually improving their performance at the For many applications of reinforcement learn- task as learning progresses. Recent work in deep RL as- ing it can be more convenient to specify both sumes that agen...