2025
Probabilistic Shielding for Safe Reinforcement Learning
at least some probability p. This framework comes from a probabilistic version of what is usually called the “safety” In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximise their reward, must often also behave in a fragment of Linear Temporal Logic (LTL) (...