Reinforcement Learning in Robotics Open access

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

Aniruddha Joshi, Niklas Lauffer, Sanjit Seshia

arXiv (Cornell University) | Jun 24, 2026

Abstract

Abstract

Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently addresses by aggregating rewards into a single scalar signal. While effective for simple tasks, this approach often fails to capture the full spectrum of optimal trade-offs, known as the Pareto frontier. In this paper, we introduce a novel preference-conditioned Bellman operator, motivated from the Chebyshev scalarization, designed to compute deterministic Pareto-optimal policies for Multi-Objective Markov Decision Processes (MOMDPs). We prove that this operator satisfies an enveloping property, where the estimated value functions upper-bound the true Pareto frontier, and demonstrate that it monotonically converges to a coverage set of this frontier. Furthermore, we also show how to extract deterministic policies from these converged Q-estimates. This ensures the agent can recover a policy for any given preference, capturing the entire Pareto-optimal frontier while guaranteeing each synthesized policy remains approximately Pareto-optimal. Experimental results validate that our algorithm successfully recovers complex trade-offs, providing a solution for deterministic Pareto-optimal policy synthesis.

Direct answer

What can I do from this paper page?

Use this page to scan "Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Reinforcement Learning in Robotics research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Aniruddha Joshi

first

Niklas Lauffer

middle

Sanjit Seshia

last

Research areas

Follow related topics

Citation

BibTeX

@article{Joshi2026Deterministic,
  title = {Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning},
  author = {Aniruddha Joshi and Niklas Lauffer and Sanjit Seshia},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.26397},
  url = {https://doi.org/10.48550/arxiv.2606.26397}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Reinforcement Learning in Robotics research papers?

Follow Reinforcement Learning in Robotics research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app