2023

Disentangling Interaction using Maximum Entropy Reinforcement Learning in Multi-Agent Systems

Paper page PDF DOI
Year
2023
Venue
ECAI 2023

Abstract

. Research on multi-agent interaction involving both mul- tiple artificial agents and humans is still in its infancy. Most recent ap- proaches have focused on environments with collaboration-focused human behavior, or providing only a small, defined set of situations. When deploying robots in human-inhabited environments in the fu- ture, it will be unlikely that all interactions fit a predefined model of collaboration, where collaborative behavior is still expected from the robot. Existing approaches are unlikely to effectively create such be- haviors in such "coexistence" environments. To tackle this issue, we introduce a novel framework that decomposes interaction and task- solving into separate learning problems and blends the resulting poli- cies at inference time. Policies are learned with maximum entropy re- inforcement learning, allowing us to create interaction-impact-aware agents and scale the cost of training agents linearly with the number Figure 1: Classification of multi-agent environment structures. Co- of agents and available tasks. We propose a weighting function cov- existence environments constrain performance of multiple agents to be interdependent but not exclusive. Cooperative environments and ering the alignment of interaction distributions with the original task. some competitive environments are considered sub-sets. We demonstrate that our framework addresses the scaling problem while solving a given task and considering collaboration opportuni- ties in a co-existence particle environment and a new cooking envi- person starts cooking something for herself. The robot and the cook ronment. Our work introduces a new learning paradigm that opens both require access to some shared resources, like the fridge, while the path to more complex multi-robot, multi-human interactions. they do not share a common goal. However, a robot that does not acknowledge the right of a bystander to approach its own goals will likely not be accepted by society but