Abstract
Abstract
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition assumptions, or alternating best-response updates, which can converge to suboptimal Nash Equilibria. In this paper, we show the joint policy gradient admits an exact decentralized decomposition of per-agent terms, each formed from per-agent score functions and decentralized critics. Based on this decomposition, we develop Agent-Chained Policy Optimization (ACPO), where actors are trained independently, with their updates together constituting a single step on the joint policy gradient. Central to this result is a serialized view of the simultaneous joint decision in which agents commit actions one at a time, each conditioning on a belief over preceding actions. The belief acts as the coordination mechanism which ties the independent per-agent updates into a joint gradient step. We evaluate ACPO on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo, where it outperforms strong baselines, with the gap widening as the number of agents grows.
Direct answer
What can I do from this paper page?
Use this page to scan "ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Reinforcement Learning in Robotics research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Matsunaga2026ACPO,
title = {ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning},
author = {Daiki E. Matsunaga and Junho Na and Tri Wahyu Guntara and Scott Sanner and Pascal Poupart and Jongmin Lee and Kee-Eung Kim},
journal = {arXiv (Cornell University)},
year = {2026},
doi = {10.48550/arxiv.2606.30072},
url = {https://doi.org/10.48550/arxiv.2606.30072}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Reinforcement Learning in Robotics research papers?
Follow Reinforcement Learning in Robotics research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app