2018

Safe Reinforcement Learning via Shielding

Paper page PDF DOI
Year
2018
Authors
Mohammed Alshiekh, Roderick Bloem, Ruediger Ehlers, Bettina Koenighofer, Scott Niekum, Ufuk Topcu
Venue
AAAI 2018
Keywords
Machine Learning Methods Track

Abstract

reward Environment Learning Agent Reinforcement learning algorithms discover policies that observation maximize reward, but do not necessarily guarantee safety dur- actions ing learning or execution phases. We introduce a new ap- proach to learn optimal policies while enforcing properties Shield expressed in temporal logic. To this end, given the temporal safe action logic specification that is to be obeyed by the learning system, we propose to synthesize a reactive system called a shield. The shield monitors the actions from the learner and corrects Figure 1: Shielded reinforcement learning them only if the chosen action causes a violation of the spec- ification. We discuss which requirements a shield must meet to preserve the convergence guarantees of the learner. Finally, its operation whenever absolutely needed in order to ensure we demonstrate the versatility of our approach on several safety?” challenging reinforcement learning scenarios. In this paper, we introduce shielded learning, a frame- work that allows applying machine learning to control sys-