Scollr summary
What this paper is about
This work presents a hybrid multimodal captioning methodology that tightly couples semantic segmentation outputs (via a LoRA-adapted Segment Anything Model) with a small, high-quality LLM- Mistral to produce descriptive, interpretable, and data-grounded scene captions.
Full abstract
Read the full abstract
Abstract. In the era of explainable AI, rapid data processing, analysis, and generation have become essential. Over the past few years, many approaches have been developed to process such heavy data and present it in an explainable manner, including in the field of remote sensing. One of such applications is remote sensing scene description. Many established workflows and models exist, but these models either fail to incorporate essential geospatial information or suffer from hallucination. We present a hybrid multimodal captioning methodology that tightly couples semantic segmentation outputs (via a LoRA-adapted Segment Anything Model) with a small, high-quality LLM- Mistral to produce descriptive, interpretable, and data-grounded scene captions. Rather than relying on direct image-to-text pipelines, our approach first extracts structured scene statistics (class proportions), spatial context (quadrant dominance and object localization), and color fingerprints (dominant colors per semantic class). These structured signals are converted into compact, factual prompts that the LLM consumes to generate coherent, informative, and verifiable captions. A comparison with the established Florence-2 model in terms of quantitative description demonstrates a significant improvement, with the Precision Vocabulary Index increasing from 0.077 to 0.232 due to the proposed workflow.
Direct answer
What can I do from this paper page?
Use this page to scan "Segmentation-driven statistics-aware workflow for detailed scene description of UAV images using Mistral and LORA powered model" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Multimodal Machine Learning Applications research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Parulekar2026Segmentation,
title = {Segmentation-driven statistics-aware workflow for detailed scene description of UAV images using Mistral and LORA powered model},
author = {Bhargav Parulekar and Anandakumar M. Ramiya},
journal = {ISPRS annals of the photogrammetry, remote sensing and spatial information sciences},
year = {2026},
doi = {10.5194/isprs-annals-xi-3-2026-231-2026},
url = {https://doi.org/10.5194/isprs-annals-xi-3-2026-231-2026}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Multimodal Machine Learning Applications research papers?
Follow Multimodal Machine Learning Applications research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Segmentation-driven statistics-aware workflow for detailed scene description of UAV images using Mistral and LORA powered model. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app