Advanced Image and Video Retrieval Techniques Open access

MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors

Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev, German Devchich and 1 more

arXiv (Cornell University) | Jul 20, 2026

Abstract

Abstract

Classical image correspondence is solved at the level of sparse keypoints or dense pixels, but the systems that consume these matches - object-level mapping, topological navigation, scene-graph maintenance - reason about whole objects. Recent work narrows this gap by matchng directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are obtained by pooling features from large 3D foundation models over the masks. We build on this segment-level matching paradigm and propose three learned matching heads: a LightGlue-style attention head with DoubleSoftmax scoring on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from the VGGT foundation model before pooling; and - as our main contribution - a multi-view extension that performs joint self-attention over segments drawn from several views at once, recovering transitive correspondences that strictly pairwise matchers cannot reach. Under a stratified zero-shot protocol on Replica and Virtual KITTI 2 with controlled viewpoint baselines from 0 deg to 180 deg, the LightGlue-style head improves over a parameter-free Sinkhorn matcher on the same MASt3R backbone by +4.85 AUPRC on Replica and +25.9 AUPRC on Virtual KITTI 2. Dropped into the RoboHop topological navigation pipeline on the Habitat-Matterport 3D (HM3D) Instance Image Navigation benchmark without retraining, our multi-view variant raises success rate from 50% to 70%, and our LightGlue-style head raises SPL from 45.7 to 59.1.

Direct answer

What can I do from this paper page?

Use this page to scan "MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Advanced Image and Video Retrieval Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Denis Fatykhoph

first

Timur Akhtyamov

middle | ORCID 0000-0003-0877-3318

Konstantin Pakulev

middle | ORCID 0000-0001-7627-0400

German Devchich

middle

Gonzalo Ferrer

last

Research areas

Follow related topics

Citation

BibTeX

@article{Fatykhoph2026MuViSeg,
  title = {MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors},
  author = {Denis Fatykhoph and Timur Akhtyamov and Konstantin Pakulev and German Devchich and Gonzalo Ferrer},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.17938},
  url = {https://doi.org/10.48550/arxiv.2607.17938}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Advanced Image and Video Retrieval Techniques research papers?

Follow Advanced Image and Video Retrieval Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app