Multimodal Machine Learning Applications Open access Peer reviewed

Object Embedding-Based Knowledge Distillation for Enhanced Visual Question Answering

Himel Das Gupta, Victor S. Sheng

Neural Processing Letters | Jul 11, 2026

Abstract

Abstract

With the rapid advancement of AI, multi-modal tasks have become key components in enhancing machine intelligence. We can observe their presence in everyday technology. A prominent example is Visual Question Answering (VQA), where users interact with AI systems to receive contextual responses based on visual input. Applications such as real-time translation, image captioning, and information retrieval through smartphone cameras highlight the growing impact of multi-modal AI in our daily lives. While these innovations make AI indispensable for modern problem-solving, their implementation often requires significant computational resources. To address this, model compression techniques such as Knowledge Distillation (KD) have been highly effective. KD aims to improve the performance of compact models by transferring knowledge from larger, more cumbersome models. However, traditional KD methods typically transfer only the probability distribution of the final layer, which may not capture the full complexity of the knowledge in tasks like VQA. In this paper, we propose a novel approach where, instead of transferring probabilities, we use object embeddings as the source of rich knowledge. These embeddings, learned from a high-performing teacher model, provide a deeper level of knowledge transfer. This rich presentation enables the compact model to better understand visual and contextual relationships, ultimately improving its performance on complex VQA tasks.

Direct answer

What can I do from this paper page?

Use this page to scan "Object Embedding-Based Knowledge Distillation for Enhanced Visual Question Answering" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Himel Das Gupta

first | Louisiana State University of Alexandria | ORCID 0000-0002-8087-9339

Victor S. Sheng

last | Texas Tech University

Research areas

Follow related topics

Citation

BibTeX

@article{Gupta2026Object,
  title = {Object Embedding-Based Knowledge Distillation for Enhanced Visual Question Answering},
  author = {Himel Das Gupta and Victor S. Sheng},
  journal = {Neural Processing Letters},
  year = {2026},
  doi = {10.1007/s11063-026-11871-0},
  url = {https://doi.org/10.1007/s11063-026-11871-0}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Object Embedding-Based Knowledge Distillation for Enhanced Visual Question Answering. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app