Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
Researchers have introduced SMART-HC-VQA, a novel dataset designed for the spatiotemporal analysis of human activity using Sentinel-2 satellite imagery. Derived from the IARPA SMART Heavy Construction dataset, this resource transforms construction-site annotations, geographic metadata, and temporal phases into natural language question-answer triplets. The dataset includes over 21,000 image chips, 65,000 single-image visual question answering (VQA) examples, and approximately 2.3 million two-image temporal comparison examples generated through Image-Pairwise Combinatorial Augmentation. Additionally, the study presents a multi-image Multimodal Large Language Model (MLLM) training framework based on LLaVA-NeXT Mistral-7B, adapted to process multiple dated image inputs. This approach redefines remote sensing challenges by treating fixed geospatial sites as evolving targets, enabling systems to not only detect changes but also reason about ongoing processes and future developments. The work provides a reproducible foundation for language-guided remote sensing activities, significantly advancing the capability of AI models to interpret complex geospatial-temporal data in construction and heavy industry contexts.
Wire timeline
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
Researchers have introduced SMART-HC-VQA, a novel dataset designed for the spatiotemporal analysis of human activity using Sentinel-2 satellite imagery. Derived from the IARPA SMART Heavy Construction dataset, this resource transforms construction-site annotations, geographic metadata, and temporal phases into natural language question-answer triplets. The dataset includes over 21,000 image chips, 65,000 single-image visual question answering (VQA) examples, and approximately 2.3 million two-image temporal comparison examples generated through Image-Pairwise Combinatorial Augmentation. Additionally, the study presents a multi-image Multimodal Large Language Model (MLLM) training framework based on LLaVA-NeXT Mistral-7B, adapted to process multiple dated image inputs. This approach redefines remote sensing challenges by treating fixed geospatial sites as evolving targets, enabling systems to not only detect changes but also reason about ongoing processes and future developments. The work provides a reproducible foundation for language-guided remote sensing activities, significantly advancing the capability of AI models to interpret complex geospatial-temporal data in construction and heavy industry contexts.
cs.AI updates on arXiv.org