Multimodal Machine Learning Applications Open access

Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation

Shuitsu Koyama, Kazuki Matsuda, Yuiga Wada, S. Hirano and 2 more

arXiv (Cornell University) | Jun 29, 2026

Abstract

Abstract

Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgments. Recent approaches using large language models (LLMs), commonly referred to as LLM-as-a-Judge, have improved alignment with human judgments but still suffer from a mismatch between large-vocabulary language modeling and evaluation over a small label set. To address this, we propose Rigel, an automatic evaluation metric for image and video captioning, based on self-distilled score adaptation. The metric employs an evaluation-specific scoring head distilled from a frozen LLM, which captures judgment signals in a task-aligned space without relying on large-vocabulary token sets. We then refine the LLM backbone with human judgment data. To train Rigel, we constructed the Vid-Lepus dataset, which contains 3,338 video clips, 33,380 reference captions, and 5,637 candidate captions. Experiments on multiple benchmarks show that Rigel outperforms state-of-the-art metrics, achieving over 10-point improvements on ActivityNet-Fact in the reference-free setting.

Direct answer

What can I do from this paper page?

Use this page to scan "Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Shuitsu Koyama

first

Kazuki Matsuda

middle

Yuiga Wada

middle | ORCID 0000-0003-3804-4546

S. Hirano

middle

Daichi Yashima

middle

Komei Sugiura

last | ORCID 0000-0002-0261-0510

Research areas

Follow related topics

Citation

BibTeX

@article{Koyama2026Rigel,
  title = {Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation},
  author = {Shuitsu Koyama and Kazuki Matsuda and Yuiga Wada and S. Hirano and Daichi Yashima and Komei Sugiura},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.29997},
  url = {https://doi.org/10.48550/arxiv.2606.29997}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app