4 papers
ai-safety
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
approaches aim to provide high-assurance quantitative guar- Ensuring that AI systems reliably and robustly antees about the safety of an AI system’s behaviour through avoid harmful or dangerous behaviours is a cru- the use of three core components — a formal safety specifi- cial...
An Algebraic Exposition of the Theory of Dyadic Morality
to be pro-social (Abdulhai et al. 2024). These different ap- arXiv:2605.16153v1 [cs.AI] 15 May 2026 proaches can be operationalized in guiding the behaviors of This paper provides an algebraic exposition of the theory AI agents with varying ease due to their inherent underly- of...
The Multi-Agent Off-Switch Game
the waiting agent off; if the agent isn’t turned off, it can proceed The off-switch game framework has been instrumental in under- to take the action. This elegant result suggests that uncertainty standing corrigibility — the property that AI agents should allow about human prefe...