Scollr summary
What this paper is about
Experiments show that Q-GrAM achieves strong text-to-image retrieval performance against both global embedding baselines and representative fine-grained image–text matching methods, while maintaining competitive bidirectional retrieval performance.
Full abstract
Read the full abstract
Existing image–text retrieval methods often compute cross-modal similarity using global single-vector representations. Although efficient for coarse semantic alignment, such compressed representations are limited when textual queries involve fine-grained semantics, including objects, attributes, relations, and their compositional structures. This paper focuses on fine-grained text-to-image retrieval and proposes Q-GrAM, a retrieval-oriented adaptation of the BLIP-2 Q-Former. Instead of treating Q-Former queries as a homogeneous set, Q-GrAM partitions a fixed query budget into semantically differentiated groups. A text-guided router assigns token-level semantic demands to query groups, while query conditional initialization modulates each group according to group-level textual summaries. The resulting grouped visual query features are matched with text tokens through a group-aware late interaction scorer, and auxiliary routing balance and inter-group diversity regularization are introduced to stabilize semantic specialization. Experiments on MS-COCO 5K, Flickr30K, and Flickr30K-CFQ show that Q-GrAM achieves strong text-to-image retrieval performance against both global embedding baselines and representative fine-grained image–text matching methods, while maintaining competitive bidirectional retrieval performance. These results demonstrate the effectiveness of structured, text-conditioned Q-Former query specialization for fine-grained text-driven image search.
Direct answer
What can I do from this paper page?
Use this page to scan "Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Gu2026GrAM,
title = {Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation},
author = {Guihe Gu and Hui Li and Hong Qin},
journal = {Sensors},
year = {2026},
doi = {10.3390/s26134313},
url = {https://doi.org/10.3390/s26134313}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?
Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app