Advanced Image and Video Retrieval Techniques Open access

Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval

R K Li, Xiaoxu Ma, Zhenyu Weng, Yue Zhang and 4 more

arXiv (Cornell University) | Jul 1, 2026

Abstract

Abstract

Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual semantic annotation. However, existing unsupervised methods rely heavily on large-scale image-text pairs. Collecting such data can be costly, particularly in scenarios where well-aligned pairs are scarce due to privacy and specialized constraints. More critically, existing methods tend to overfit to seen training data, restricting their generalization performance on unseen categories that the constrained training data cannot cover. To address these limitations, we propose Attribute-Prompted Kernel Hashing (APKH), a novel data-efficient approach that constructs a compact, modality-aligned Hamming space driven by the generalized attribute priors of vision-language foundation models. Specifically, APKH introduces two core modules: Context-optimized Attribute Kernel Mapping (CAKM) and Kernel-Smoothed Contrastive Alignment (KSCA). CAKM formulates cross-modal alignment through hyperspherical Radial Basis Function kernel mapping, optimizing dynamic attribute kernels via prompt learning to capture modality-invariant semantics. Furthermore, KSCA extends conventional point-to-point contrastive learning by modeling limited paired data as continuous kernel distributions. This explicit smoothing of the modality gap alleviates overfitting to sparse pairwise correlations. Extensive experiments demonstrate that APKH outperforms state-of-the-art hashing methods in the challenging cross-modal retrieval tasks from seen to unseen categories under data-constrained scenarios.

Direct answer

What can I do from this paper page?

Use this page to scan "Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

R K Li

first

Xiaoxu Ma

middle

Zhenyu Weng

middle | ORCID 0000-0001-7857-8687

Yue Zhang

middle

Guibo Luo

middle

Huiping Zhuang

middle

Zhiping Lin

middle

Yap-Peng Tan

last

Research areas

Follow related topics

Citation

BibTeX

@article{Li2026Attribute,
  title = {Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval},
  author = {R K Li and Xiaoxu Ma and Zhenyu Weng and Yue Zhang and Guibo Luo and Huiping Zhuang and Zhiping Lin and Yap-Peng Tan},
  journal = {arXiv (Cornell University)},
  year = {2026},
  url = {https://arxiv.org/abs/2607.00379}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?

Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app