2024

Safe Model-Based Multi-Agent Mean-Field Reinforcement Learning

Paper page PDF
Year
2024
arXiv
2306.17052 [cs.LG]

Abstract

Many applications, e.g., in shared mobility, require coordinating a large number of agents. Mean-field reinforcement learning ad- dresses the resulting scalability challenge by optimizing the policy of a representative agent interacting with the infinite population of identical agents instead of considering individual pairwise interac- tions. In this paper, we address an important generalization where there exist global constraints on the distribution of agents (e.g., re- quiring capacity constraints or minimum coverage requirements to be met). We propose Safe-M3 -UCRL, the first model-based mean- field reinforcement learning algorithm that attains safe policies even in the case of unknown transitions. As a key ingredient, it uses epistemic uncertainty in the transition model within a log-barrier approach to ensure pessimistic constraints satisfaction with high probability. Beyond the synthetic swarm motion benchmark, we showcase Safe-M3 -UCRL on the vehicle repositioning problem faced by many shared mobility operators and evaluate its perfor- mance through simulations built on vehicle trajectory data from a service provider in Shenzhen. Our algorithm effectively meets Figure 1: An illustration of vehicles’ spatial distribution the demand in critical areas while ensuring service accessibility in (light-blue scatters), repositioning trips (blue arrows), and a regions with low demand. trajectory of passenger trips (red arrows).