Reinforcement Learning in Robotics Open access

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Lars van der Laan, Nathan Kallus

arXiv (Cornell University) | Jul 6, 2026

Abstract

Abstract

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback--Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to log-ratio approximation error and a statistical error governed by the complexity of the ratio hypothesis class. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.

Direct answer

What can I do from this paper page?

Use this page to scan "Fitted Occupancy-Ratio Evaluation without Bellman Completeness" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Reinforcement Learning in Robotics research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Lars van der Laan

first

Nathan Kallus

last

Research areas

Follow related topics

Citation

BibTeX

@article{Laan2026Fitted,
  title = {Fitted Occupancy-Ratio Evaluation without Bellman Completeness},
  author = {Lars van der Laan and Nathan Kallus},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.05375},
  url = {https://doi.org/10.48550/arxiv.2607.05375}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Reinforcement Learning in Robotics research papers?

Follow Reinforcement Learning in Robotics research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Fitted Occupancy-Ratio Evaluation without Bellman Completeness. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app