Multimodal Machine Learning Applications Open access

Evidence-Backed Video Question Answering

Shijie Wang, Honglu Zhou, Ziyang Wang, Ran Xu and 4 more

arXiv (Cornell University) | Jul 13, 2026

Abstract

Abstract

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. To support this, we introduce ST-Evidence, the first human-verified benchmark for both discriminative and generative pixel-level grounding. Evaluations of state-of-the-art models reveal a critical decoupling between QA accuracy and true visual perception that scaling alone fails to bridge. To address this, we develop scalable, automated generation pipelines to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding. Fine-tuning grounded Video LLMs on this data yields substantial gains over the corresponding size-matched UniPixel baselines (e.g., +27.2 t-mean and +13.8 J&F on a 7B model), establishing a robust baseline for explainable, evidence-backed video understanding. Code and data are available at https://github.com/SalesforceAIResearch/EVQA.

Direct answer

What can I do from this paper page?

Use this page to scan "Evidence-Backed Video Question Answering" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Shijie Wang

first

Honglu Zhou

middle

Ziyang Wang

middle

Ran Xu

middle

Caiming Xiong

middle

Silvio Savarese

middle

Chen Sun

middle

Juan Carlos Niebles

last

Research areas

Follow related topics

Citation

BibTeX

@article{Wang2026Evidence,
  title = {Evidence-Backed Video Question Answering},
  author = {Shijie Wang and Honglu Zhou and Ziyang Wang and Ran Xu and Caiming Xiong and Silvio Savarese and Chen Sun and Juan Carlos Niebles},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.11862},
  url = {https://doi.org/10.48550/arxiv.2607.11862}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Evidence-Backed Video Question Answering. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app