Multimodal Machine Learning Applications Open access

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo, Angélique Loesch

arXiv (Cornell University) | Jun 22, 2026

Abstract

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. Grounded Visual Reasoning (GVR) approaches attempt to improve interpretability by explicitly couple textual rationales with visual grounding information, which are typically textual coordinates. This mechanism lacks a learnable semantic link to the visual features, often resulting in a semantic-spatial gap where the model hallucinates coordinates that do not correspond to image evidences. In this work, we introduce Composer, a MLLM that leverages a novel visual grounding mechanism based on learned proxy-tokens to promote faithful interpretability. These discrete symbolic pointers explicitly index the image latent space, allowing the model to manipulate visual regions as addressable, semantically manipulable sets. To rigorously validate our novel grounding mechanism, we constructed ComposerGCoT, a dataset synthesized to enable holistic assessment of reasoning consistency and grounding accuracy. Experimental results indicate that Composer achieves performance parity with its coordinate-based counterpart in final answer accuracy, while improving visual grounding accuracy by +9.0 points. By demonstrating that discrete proxy-tokens capture spatial semantics more effectively than typical textual coordinates, we establish that visual grounding mechanisms with learnable semantic links represent a promising path toward trustworthy and reliable MLLMs.

Direct answer

What can I do from this paper page?

Use this page to scan "Faithful Grounded Visual Reasoning via Learned Proxy-Tokens" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Tom Hodemon

first

Mohamed Chaouch

middle

Aboubacar Tuo

middle | ORCID 0000-0002-4045-8660

Angélique Loesch

last | ORCID 0000-0001-5427-3010

Research areas

Follow related topics

Citation

BibTeX

@article{Hodemon2026Faithful,
  title = {Faithful Grounded Visual Reasoning via Learned Proxy-Tokens},
  author = {Tom Hodemon and Mohamed Chaouch and Aboubacar Tuo and Angélique Loesch},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.23354},
  url = {https://doi.org/10.48550/arxiv.2606.23354}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Faithful Grounded Visual Reasoning via Learned Proxy-Tokens. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app