Multi-Agent Reinforcement Learning: Cooperative, Competitive, and Mixed-Objective Methods
Summary
This introductory survey organizes multi-agent reinforcement learning (MARL) by the relationships among agents’ objectives: fully cooperative, fully competitive, and mixed cooperation and competition. It explains that cooperative agents can be modeled jointly by treating their actions as one vector, or separately through independent learners. Independent Q-learning is simple, but agents’ interacting actions can make coordination difficult and prevent convergence to a globally effective policy.
For competitive settings, the article introduces minimax reasoning and minimax Q-learning. For mixed settings, a prisoner’s dilemma example illustrates how individually rational choices can lead to a worse joint outcome than cooperation. It names Nash-Q as a cautious equilibrium approach and lists WoLF-IGA, WoLF-PHC, GIGA, GIGA-WoLF, and CE-Q as methods aimed at adapting behavior or supporting cooperation. The discussion is conceptual: it gives no formal algorithm derivations, comparative experiments, or trading applications, so it serves as orientation rather than implementation guidance.
Key ideas
- MARL problems can be grouped by whether agents share rewards, oppose one another, or have mixed incentives.
- Cooperative agents can be modeled as one joint policy or as independent learners, though independent learning may fail to coordinate.
- Minimax Q-learning extends adversarial minimax reasoning to competitive multi-agent learning.
- In mixed-objective games, individually rational choices can produce worse collective rewards than cooperation.
- The article names equilibrium and adaptive-learning algorithms but does not derive or experimentally compare them.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.