Abstract
Abstract
Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on DocVQA in three different document domains: industrial documents of varying type, infographics, and presentation slides. We systematically assess model performance under zero-shot evaluations, fully supervised finetuning with inter- and intra-dataset evaluations, and few-shot learning evaluations of knowledge transfer between domains. Our findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides. Although parameter scaling is a dominant factor on performance, supervised finetuning yields higher relative gains in smaller architectures. Furthermore, our cross-domain and few-shot experiments show that visual understanding is the main bottleneck for DocVQA, not a lack of knowledge from the VLMs. Using 50 target domain samples, the models finetuned in DocVQA with datasets of different domains rapidly adapt to the target domain documents, even surpassing their fully supervised counterparts in some cases.
Direct answer
What can I do from this paper page?
Use this page to scan "Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{LopezDuran2026Comparative,
title = {Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering},
author = {Miguel Lopez-Duran and Elena Marrero and Julian Fierrez and Marta Robledo-Moreno and Ruben Vera-Rodriguez and Daniel DeAlcala and Aythami Morales and Ruben Tolosana and Oscar Delgado and Álvaro Ortigosa and Javier Ortega-García},
journal = {arXiv (Cornell University)},
year = {2026},
doi = {10.48550/arxiv.2607.07179},
url = {https://doi.org/10.48550/arxiv.2607.07179}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Multimodal Machine Learning Applications research papers?
Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app