Abstract
Abstract
Indoor visual place recognition (VPR) is challenged by repetitive architec tural layouts where appearancebased retrieval confuses visually similar floors and corridors.This limitation reduces reliability for robotic navigation in large indoor buildings, including practical deployments in Vietnam.Scene text provides high-level semantic cues, yet prior text-aware re-ranking commonly relies on non-learnable string overlap and remains sensitive to optical character recognition (OCR) noise while exploiting spatial visual and textual interactions only weakly.TiP-MCE is introduced as an enhancement over TextInPlace by replacing heuristic text matching with a learnable multimodal cross-encoder (MCE) re-ranker.For each query and each candidate in the top-K list, visual tokens and filtered text instances with content embeddings, bounding boxes, and confidence scores are jointly processed by Transformer self attention to estimate the probability that both images depict the same place.A two-phase training protocol is adopted.Metric learning is used to train the global retrieval branch, then MCE is trained with binary crossentropy on hard negatives mined from top-M retrieval results using pose distance supervision.Experiments on Maze-with-Text show the best overall Recall@1 of 86.8%, surpassing TextInPlace at 85.4% and TextInPlace-I at 86.1%, with larger gains on harder floors and competitive Recall@5.
Direct answer
What can I do from this paper page?
Use this page to scan "TiP-MCE: Multimodal Cross-encoder Re-ranking for Indoor Visual Place Recognition" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Nguyen2026Multimodal,
title = {TiP-MCE: Multimodal Cross-encoder Re-ranking for Indoor Visual Place Recognition},
author = {Khac Toan Nguyen and Bao T. Nguyen and Tuyet Hue Tran},
journal = {International journal of intelligent engineering and systems},
year = {2026},
doi = {10.22266/ijies2026.0731.44},
url = {https://doi.org/10.22266/ijies2026.0731.44}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?
Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for TiP-MCE: Multimodal Cross-encoder Re-ranking for Indoor Visual Place Recognition. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app