Reinforcement Learning in Robotics Open access

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

Pascal Wilhelm, Odej Kao

arXiv (Cornell University) | Jul 15, 2026

Abstract

Abstract

Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget. We study the fixed-budget decision problem behind this practice: under the same post-training budget, should one use a larger policy, train a smaller policy longer, generate more rollout search, or spend compute on stronger reward feedback? We introduce a FLOP-accounting framework for GRPO post-training that decomposes compute into rollout/search, policy-update/learning, and reward- or feedback-model evaluation. Across LoRA-adapted Qwen2.5 policies, we find conditional allocation frontiers: the best observed allocation changes with model size, compute budget, reward system, and evaluation target. Same-FLOP model-size comparisons show that model choice and training allocation are coupled because larger policies consume more per-token compute and therefore buy fewer updates or rollouts under the same budget. Reward systems also change the accounting: rule-based rewards spend nearly all non-update compute on policy rollouts, while PRM-style feedback allocates a visible part of the budget to reward-model inference. We present RACE as a diagnostic pilot-grid protocol, not a guarantee of held-out improvement, for identifying allocation regimes before expensive validation runs; our results suggest that RL post-training papers should report total FLOPs together with how compute is divided among model size, search, learning, and feedback.

Direct answer

What can I do from this paper page?

Use this page to scan "Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Reinforcement Learning in Robotics research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Pascal Wilhelm

first

Odej Kao

last

Research areas

Follow related topics

Citation

BibTeX

@article{Wilhelm2026Where,
  title = {Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback},
  author = {Pascal Wilhelm and Odej Kao},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.13389},
  url = {https://doi.org/10.48550/arxiv.2607.13389}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Reinforcement Learning in Robotics research papers?

Follow Reinforcement Learning in Robotics research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app