2018
Safe Reinforcement Learning via Shielding
- Year
- 2018
- Authors
- Mohammed Alshiekh, Roderick Bloem, Ruediger Ehlers, Bettina Koenighofer, Scott Niekum, Ufuk Topcu
- Venue
- AAAI 2018
- Keywords
- Machine Learning Methods Track
Abstract
reward Environment Learning Agent Reinforcement learning algorithms discover policies that observation maximize reward, but do not necessarily guarantee safety dur- actions ing learning or execution phases. We introduce a new ap- proach to learn optimal policies while enforcing properties Shield expressed in temporal logic. To this end, given the temporal safe action logic specification that is to be obeyed by the learning system, we propose to synthesize a reactive system called a shield. The shield monitors the actions from the learner and corrects Figure 1: Shielded reinforcement learning them only if the chosen action causes a violation of the spec- ification. We discuss which requirements a shield must meet to preserve the convergence guarantees of the learner. Finally, its operation whenever absolutely needed in order to ensure we demonstrate the versatility of our approach on several safety?” challenging reinforcement learning scenarios. In this paper, we introduce shielded learning, a frame- work that allows applying machine learning to control sys-