2024

Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

Paper page PDF
Year
2024
Authors
David ``davidad'' Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, Joshua Tenenbaum
arXiv
2405.06624 [cs.AI]
Keywords
Machine Learning, ICML

Abstract

approaches aim to provide high-assurance quantitative guar- Ensuring that AI systems reliably and robustly antees about the safety of an AI system’s behaviour through avoid harmful or dangerous behaviours is a cru- the use of three core components — a formal safety specifi- cial challenge, especially for AI systems with a cation, a world model, and a verifier. We will argue that this high degree of autonomy and general intelligence, strategy is both promising and underexplored, and contrast it or systems used in safety-critical contexts. In this with other ongoing efforts in AI safety. We will also outline paper, we will introduce and define a family of several ongoing avenues of research within the broader GS approaches to AI safety, which we will refer to research agenda, identify some of their core difficulties, and as guaranteed safe (GS) AI. The core feature of discuss approaches for overcoming these difficulties. Cen- these approaches is that they aim to produce AI tral examples of agendas which fall under the GS AI family systems which are equipped with high-assurance include Szegedy (2020); Wing (2021); Seshia et al. (2022); quantitative safety guarantees. This is achieved Russell (2022); Tegmark & Omohundro (2023); ’davidad’ by the interplay of three core components: a world Dalrymple (2024); Bengio (2024). model (which provides a mathematical descrip- Critical infrastructure and safety-critical systems are re- tion of how the AI system affects the outside quired to comply with high safety standards. For example, world), a safety specification (which is a math- aircrafts, nuclear power plants, and medical devices are sub- ematical description of what effects are accept- ject to exceptionally rigorous safety certification. Moreover, able), and a verifier (which provides an auditable it is plausible that there will soon be AI systems that are proof certificate that the AI satisfies the safety at least as safety-critical as these systems (e.g