Multimodal Machine Learning Applications Open access

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo and 1 more

arXiv (Cornell University) | Jun 28, 2026

Abstract

Abstract

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. However, existing benchmarks primarily evaluate whether models can perceive shallow visual cues, while rarely examining whether MLLMs can learn deeper knowledge or procedural skills from video tutorials and generalize them to downstream long-horizon agentic tasks. To address this gap, we introduce VG-GUIBench (Video-Guided GUI Benchmark), a new benchmark designed to evaluate whether MLLM-based GUI agents can follow video tutorials to complete corresponding GUI interactive tasks. Furthermore, we observe that the performance of models on both VideoQA and video-guided agentic tasks critically depends on effective keyframe extraction. Based on this observation, we propose TASKER (Task-driven And Scene-aware Keyframe searchER), a keyframe extraction algorithm that jointly considers task relevance and scene dynamics to identify informative frames. Experimental results demonstrate that TASKER achieves significant performance improvements on both VideoQA and video-guided agentic task benchmarks, outperforming the best baseline by 2.0% on the EgoSchema fullset and 1.8% on the NExT-QA dataset, respectively. These results further highlight the potential of generalized keyframe extraction methods for video understanding tasks. Our code and data are available at https://github.com/VG-GUI-TASKER/VG-GUI-TASKER.

Direct answer

What can I do from this paper page?

Use this page to scan "Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Sunqi Fan

first

Qingle Liu

middle

Runqi Yin

middle

Meng-Hao Guo

middle | ORCID 0000-0002-4128-4594

Sen Yang

last | ORCID 0000-0003-1064-0055

Research areas

Follow related topics

Citation

BibTeX

@article{Fan2026Bridging,
  title = {Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction},
  author = {Sunqi Fan and Qingle Liu and Runqi Yin and Meng-Hao Guo and Sen Yang},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.29445},
  url = {https://doi.org/10.48550/arxiv.2606.29445}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app