Advanced Image and Video Retrieval Techniques Open access

Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing

R K Li, Xiaoxu Ma, Zhenyu Weng, Yue Zhang and 4 more

arXiv (Cornell University) | Jun 30, 2026

Abstract

Abstract

Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existing unsupervised CMH methods often rely on large-scale image-text pairs, which are costly to collect. To address this limitation, we propose Global-Neighborhood Alignment Hashing (GNAH), a novel approach that preserves the semantic structure of vision-language foundation models within a compact binary Hamming space using only a limited number of image-text pairs. Specifically, GNAH captures global structural information from the continuous latent space and transfers it into the binary Hamming space through a Prototype-Anchored Global Alignment module. In addition, GNAH extends conventional pairwise contrastive learning by modeling stochastic neighborhood relationships via a Contrastive Stochastic Neighborhood Alignment module, thereby alleviating overfitting to sparse pairwise correlations. Extensive experiments demonstrate that GNAH consistently outperforms existing unsupervised cross-modal retrieval methods under data-constrained settings, offering a practical solution for real-world CMH applications.

Direct answer

What can I do from this paper page?

Use this page to scan "Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

R K Li

first

Xiaoxu Ma

middle

Zhenyu Weng

middle | ORCID 0000-0001-7857-8687

Yue Zhang

middle

Guibo Luo

middle

Huiping Zhuang

middle

Zhiping Lin

middle

Yap-Peng Tan

last

Research areas

Follow related topics

Citation

BibTeX

@article{Li2026Unsupervised,
  title = {Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing},
  author = {R K Li and Xiaoxu Ma and Zhenyu Weng and Yue Zhang and Guibo Luo and Huiping Zhuang and Zhiping Lin and Yap-Peng Tan},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.31517},
  url = {https://doi.org/10.48550/arxiv.2606.31517}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?

Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app