Reinforcement Learning in Robotics Open access

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner and 3 more

arXiv (Cornell University) | Jun 29, 2026

Abstract

Abstract

Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition assumptions, or alternating best-response updates, which can converge to suboptimal Nash Equilibria. In this paper, we show the joint policy gradient admits an exact decentralized decomposition of per-agent terms, each formed from per-agent score functions and decentralized critics. Based on this decomposition, we develop Agent-Chained Policy Optimization (ACPO), where actors are trained independently, with their updates together constituting a single step on the joint policy gradient. Central to this result is a serialized view of the simultaneous joint decision in which agents commit actions one at a time, each conditioning on a belief over preceding actions. The belief acts as the coordination mechanism which ties the independent per-agent updates into a joint gradient step. We evaluate ACPO on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo, where it outperforms strong baselines, with the gap widening as the number of agents grows.

Direct answer

What can I do from this paper page?

Use this page to scan "ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Reinforcement Learning in Robotics research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Daiki E. Matsunaga

first

Junho Na

middle | ORCID 0009-0004-1054-1975

Tri Wahyu Guntara

middle

Scott Sanner

middle

Pascal Poupart

middle

Jongmin Lee

middle

Kee-Eung Kim

last

Research areas

Follow related topics

Citation

BibTeX

@article{Matsunaga2026ACPO,
  title = {ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning},
  author = {Daiki E. Matsunaga and Junho Na and Tri Wahyu Guntara and Scott Sanner and Pascal Poupart and Jongmin Lee and Kee-Eung Kim},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.30072},
  url = {https://doi.org/10.48550/arxiv.2606.30072}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Reinforcement Learning in Robotics research papers?

Follow Reinforcement Learning in Robotics research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app