Multimodal Machine Learning Applications Open access

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno and 7 more

arXiv (Cornell University) | Jul 8, 2026

Abstract

Abstract

Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on DocVQA in three different document domains: industrial documents of varying type, infographics, and presentation slides. We systematically assess model performance under zero-shot evaluations, fully supervised finetuning with inter- and intra-dataset evaluations, and few-shot learning evaluations of knowledge transfer between domains. Our findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides. Although parameter scaling is a dominant factor on performance, supervised finetuning yields higher relative gains in smaller architectures. Furthermore, our cross-domain and few-shot experiments show that visual understanding is the main bottleneck for DocVQA, not a lack of knowledge from the VLMs. Using 50 target domain samples, the models finetuned in DocVQA with datasets of different domains rapidly adapt to the target domain documents, even surpassing their fully supervised counterparts in some cases.

Direct answer

What can I do from this paper page?

Use this page to scan "Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Miguel Lopez-Duran

first | ORCID 0009-0001-0976-7174

Elena Marrero

middle

Julian Fierrez

middle

Marta Robledo-Moreno

middle

Ruben Vera-Rodriguez

middle

Daniel DeAlcala

middle | ORCID 0000-0001-9243-6108

Aythami Morales

middle | ORCID 0000-0002-7268-4785

Ruben Tolosana

middle

Oscar Delgado

middle

Álvaro Ortigosa

middle | ORCID 0000-0002-7674-4132

Javier Ortega-García

last | ORCID 0000-0003-0557-1948

Research areas

Follow related topics

Citation

BibTeX

@article{LopezDuran2026Comparative,
  title = {Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering},
  author = {Miguel Lopez-Duran and Elena Marrero and Julian Fierrez and Marta Robledo-Moreno and Ruben Vera-Rodriguez and Daniel DeAlcala and Aythami Morales and Ruben Tolosana and Oscar Delgado and Álvaro Ortigosa and Javier Ortega-García},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.07179},
  url = {https://doi.org/10.48550/arxiv.2607.07179}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app