Advanced Image and Video Retrieval Techniques Open access Peer reviewed

Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation

Guihe Gu, Hui Li, Hong Qin

Sensors | Jul 7, 2026

Scollr summary

What this paper is about

Experiments show that Q-GrAM achieves strong text-to-image retrieval performance against both global embedding baselines and representative fine-grained image–text matching methods, while maintaining competitive bidirectional retrieval performance.

Full abstract

Read the full abstract

Existing image–text retrieval methods often compute cross-modal similarity using global single-vector representations. Although efficient for coarse semantic alignment, such compressed representations are limited when textual queries involve fine-grained semantics, including objects, attributes, relations, and their compositional structures. This paper focuses on fine-grained text-to-image retrieval and proposes Q-GrAM, a retrieval-oriented adaptation of the BLIP-2 Q-Former. Instead of treating Q-Former queries as a homogeneous set, Q-GrAM partitions a fixed query budget into semantically differentiated groups. A text-guided router assigns token-level semantic demands to query groups, while query conditional initialization modulates each group according to group-level textual summaries. The resulting grouped visual query features are matched with text tokens through a group-aware late interaction scorer, and auxiliary routing balance and inter-group diversity regularization are introduced to stabilize semantic specialization. Experiments on MS-COCO 5K, Flickr30K, and Flickr30K-CFQ show that Q-GrAM achieves strong text-to-image retrieval performance against both global embedding baselines and representative fine-grained image–text matching methods, while maintaining competitive bidirectional retrieval performance. These results demonstrate the effectiveness of structured, text-conditioned Q-Former query specialization for fine-grained text-driven image search.

Direct answer

What can I do from this paper page?

Use this page to scan "Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Guihe Gu

first | Wuhan University | ORCID 0009-0002-2456-7629

Hui Li

middle | Wuhan University | ORCID 0000-0001-9198-3951

Hong Qin

last | Wuhan University | ORCID 0000-0002-7220-5246

Research areas

Follow related topics

Citation

BibTeX

@article{Gu2026GrAM,
  title = {Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation},
  author = {Guihe Gu and Hui Li and Hong Qin},
  journal = {Sensors},
  year = {2026},
  doi = {10.3390/s26134313},
  url = {https://doi.org/10.3390/s26134313}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?

Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Q-GrAM: Fine-Grained Image–Text Retrieval via Grouped Query Routing and Conditional Query Modulation. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app