2026
The Multi-Agent Off-Switch Game
- Year
- 2026
- Venue
- AAMAS 2026
Abstract
the waiting agent off; if the agent isn’t turned off, it can proceed The off-switch game framework has been instrumental in under- to take the action. This elegant result suggests that uncertainty standing corrigibility — the property that AI agents should allow about human preferences can serve as a natural mechanism for human oversight and intervention. In single-agent settings, uncer- maintaining AI corrigibility. tainty about human preferences naturally incentivizes agents to However, modern AI systems rarely operate in isolation. From defer to human judgment. However, as AI systems increasingly autonomous trading algorithms interacting in financial markets operate in multi-agent environments, a crucial question arises: [5] to AI-powered defense systems monitoring for cyber threats does corrigibility compose across multiple agents? We introduce [21], artificial agents increasingly find themselves in multi-agent the multi-agent off-switch game and demonstrate that individu- environments where strategic considerations play a crucial role ally corrigible agents can become collectively incorrigible when in decision-making. This raises a fundamental question: does the strategic interactions are considered. Through formal analysis and corrigibility observed in single-agent settings extend to multi-agent illustrative examples, we show that corrigibility is not composi- scenarios? tional and identify conditions under which group incorrigibility In this paper, we generalize the off-switch game to the multi- emerges. Our results highlight fundamental challenges for AI safety agent case and uncover a concerning result: corrigibility is not in multi-agent settings and suggest the need for new approaches compositional. Agents that would behave corrigibly when operating that explicitly address collective dynamics. alone can become incorrigible when strategic interactions with other agents are considered. This breakdown occurs even when