Multimodal Machine Learning Applications Open access

RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory

Yixun Hu, Zhicheng Zheng, Lihan Zha, Chunwei Xing and 4 more

arXiv (Cornell University) | Jun 23, 2026

Abstract

Abstract

Long-term robot deployment requires a compact and scalable memory that preserves fine-grained visual semantics, grounds observations in space and time, and enables efficient storage and retrieval. In this paper, we propose RAVEN, an agentic memory system for long-horizon robotic question answering and navigation. RAVEN stores visual embeddings with pose and time in a vector database, and grounds retrieval in a spatial map to answer queries and navigate to goals. By operating directly on visual embeddings, RAVEN avoids lossy image-to-text captioning and enables accurate semantic, spatial, and temporal retrieval at scale. Across several simulated and real-world video question-answering benchmarks, RAVEN consistently surpasses caption-based memory systems and matches frontier VLMs on long-horizon tasks at 10$\times$ lower retrieval cost. Finally, we instantiate RAVEN on a Unitree Go1 robot for the task of long-horizon navigation for natural language goal-reaching, and show successful deployment over several large indoor environments.

Direct answer

What can I do from this paper page?

Use this page to scan "RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Yixun Hu

first

Zhicheng Zheng

middle

Lihan Zha

middle

Chunwei Xing

middle | ORCID 0000-0003-4920-3450

Rajdeep Singh

middle

Omar Hossain

middle

Antonio Loquercio

middle | ORCID 0000-0002-8410-3933

Dhruv Shah

last

Research areas

Follow related topics

Citation

BibTeX

@article{Hu2026RAVEN,
  title = {RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory},
  author = {Yixun Hu and Zhicheng Zheng and Lihan Zha and Chunwei Xing and Rajdeep Singh and Omar Hossain and Antonio Loquercio and Dhruv Shah},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2606.25206},
  url = {https://doi.org/10.48550/arxiv.2606.25206}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Multimodal Machine Learning Applications research papers?

Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app