Multimodal Machine Learning Applications Open access

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

Mingkuan Feng, Jinyang Wu, Hao Gu, Fangrui Lv and 4 more

arXiv (Cornell University) | Jun 29, 2026

Abstract

Abstract

Agentic multimodal models perform diverse operations on an image via code and reason over the returned view, an effective paradigm for fine-grained visual question answering. However, code operations can be useful, redundant, or misleading. Outcome-only rewards cannot precisely distinguish these cases, and existing process rewards either fail to attribute final correctness to individual tool calls, or require an external judge model. To address this, we introduce Tool-Augmented Credit Optimization (TACO), a GRPO variant for code-tool agents built on two coupled advantage channels. The first, Differential Answer-Probe Reward (DAPR), is a self-supervised, judge-free tool-contribution advantage that credits each tool call by its own effect on answering correctly. Probe tokens inserted into the model's reasoning elicit its predictions with and without the tool, and the difference in outcome reward is taken as the call's value: positive for a useful call, negative for a misleading one, and zero for one that changes nothing. This reuses the existing answer checker with no auxiliary judge, and, being a difference rather than an absolute probe score, is naturally robust to probe-hacking. The second is the outcome advantage from the final answer, distributed by Outcome-Gated Advantage Routing (OGAR): a parameter-free rule that, conditioned on the call's outcome, delivers this credit only to the responsible segments, suppressing wasted tool calls without any cost term. We train TACO through a two-stage SFT+RL pipeline. Extensive experiments across perception, reasoning, and general multimodal benchmarks show that it yields consistent accuracy gains and learns to invoke its tools only when they help.

Direct answer

What can I do from this paper page?

Use this page to scan "TACO: Tool-Augmented Credit Optimization for Agentic Tool Use" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Mingkuan Feng

first

Jinyang Wu

middle

Hao Gu

middle

Fangrui Lv

middle

Ruihan Jin

middle

Chuyuan Zhang

middle

Zhengqi Wen

middle

Jianhua Tao

last

Research areas

Follow related topics

Citation

BibTeX

@article{Feng2026TACO,
  title = {TACO: Tool-Augmented Credit Optimization for Agentic Tool Use},
  author = {Mingkuan Feng and Jinyang Wu and Hao Gu and Fangrui Lv and Ruihan Jin and Chuyuan Zhang and Zhengqi Wen and Jianhua Tao},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.30251},
  url = {https://doi.org/10.48550/arxiv.2606.30251}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for TACO: Tool-Augmented Credit Optimization for Agentic Tool Use. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app