このエントリーをはてなブックマークに追加
ID 70764
フルテキストURL
著者
Manatphaiboon, Natchanon Graduate School of Environmental, Life, Natural Science and Technology, Okayama University
Monden, Akito Graduate School of Environmental, Life, Natural Science and Technology, Okayama University ORCID Kaken ID researchmap
Yücel, Zeynep Department of Environmental Sciences, Informatics and Statistics, Ca’Foscari University of Venice
抄録
We study convergence in multi-agent reinforcement learning (MARL) through the lens of sufficient conditions, using a single-point-of-failure analysis applied to multi-agent policy iteration integrated with linear programming (MAPI-LP), where results are proven for pure coordination games, and extension to broader settings is conjectured. We identify two sufficient conditions for convergence to Markov Perfect Equilibrium (MPE). The first is stability in best-response space that emerges from value monotonicity. The second, monotonic best-response space shrinking (MBRSS), is a novel condition requiring that each agent’s best-response space contracts monotonically across iterations until it collapses to a stable space. Furthermore, we show that MBRSS does not necessarily imply monotonic improvement in this setting. However, value monotonicity and stability in best-response space imply each other when a complementary condition is applied. Building on this hierarchical relationship, we propose conjectures on sufficient condition relationships in both serial and parallel MARL. In addition, we propose conjectures on generalized MBRSS to arbitrary finite repeated games and validation of stability in best-response space. We further discuss connections between MBRSS and existing related frameworks, and outline directions toward a taxonomy of sufficient conditions for MARL convergence.
キーワード
sufficient condition
convergence analysis
Markov perfect equilibrium
multi-agent reinforcement learning
best-response dynamics
発行日
2026-06-15
出版物タイトル
Mathematics
14巻
12号
出版者
MDPI AG
開始ページ
2134
ISSN
2227-7390
資料タイプ
学術雑誌論文
言語
英語
OAI-PMH Set
岡山大学
著作権者
© 2026 by the authors.
論文のバージョン
publisher
DOI
関連URL
isVersionOf https://doi.org/10.3390/math14122134
ライセンス
https://creativecommons.org/licenses/by/4.0/
Citation
Manatphaiboon, N.; Monden, A.; Yücel, Z. Toward Convergence in Multi-Agent Reinforcement Learning: Best-Response Space Shrinking as a Sufficient Condition. Mathematics 2026, 14, 2134. https://doi.org/10.3390/math14122134