Abstract
Abstract
Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled image-text pairs. However, existing unsupervised CMH methods often rely on large-scale image-text pairs, which are costly to collect. To address this limitation, we propose Global-Neighborhood Alignment Hashing (GNAH), a novel approach that preserves the semantic structure of vision-language foundation models within a compact binary Hamming space using only a limited number of image-text pairs. Specifically, GNAH captures global structural information from the continuous latent space and transfers it into the binary Hamming space through a Prototype-Anchored Global Alignment module. In addition, GNAH extends conventional pairwise contrastive learning by modeling stochastic neighborhood relationships via a Contrastive Stochastic Neighborhood Alignment module, thereby alleviating overfitting to sparse pairwise correlations. Extensive experiments demonstrate that GNAH consistently outperforms existing unsupervised cross-modal retrieval methods under data-constrained settings, offering a practical solution for real-world CMH applications.
Direct answer
What can I do from this paper page?
Use this page to scan "Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Li2026Unsupervised,
title = {Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing},
author = {R K Li and Xiaoxu Ma and Zhenyu Weng and Yue Zhang and Guibo Luo and Huiping Zhuang and Zhiping Lin and Yap-Peng Tan},
journal = {arXiv (Cornell University)},
year = {2026},
doi = {10.48550/arxiv.2606.31517},
url = {https://doi.org/10.48550/arxiv.2606.31517}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?
Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app