MPerS: Dynamic MLLM MixExperts for Remote Sensing Scene Segmentation
Researchers have introduced MPerS, a novel framework for dynamic Multimodal Large Language Model (MLLM) Mixture-of-Experts perception-guided remote sensing scene segmentation. Addressing limitations in existing studies that often neglect high-quality caption generation, this method utilizes multiple prompts to enable MLLMs like LLaVA, ChatGPT, and Qwen to perceive complex remote sensing scenes from diverse expert perspectives. The approach employs DINOv3 for extracting dense visual representations of land covers and features a Dynamic MixExperts module to adaptively integrate the most effective textual semantics. Additionally, a Linguistic Query Guided Attention mechanism is constructed to leverage textual semantic information for guiding visual features, ensuring precise segmentation. The proposed method demonstrates superior performance across three public semantic segmentation remote sensing datasets, marking a significant advancement in multimodal fusion for geospatial analysis. This development highlights the growing integration of advanced AI models in interpreting complex environmental data, offering improved accuracy in scene understanding and segmentation tasks within the field of computer vision and remote sensing.
Wire timeline
MPerS: Dynamic MLLM MixExperts for Remote Sensing Scene Segmentation
Researchers have introduced MPerS, a novel framework for dynamic Multimodal Large Language Model (MLLM) Mixture-of-Experts perception-guided remote sensing scene segmentation. Addressing limitations in existing studies that often neglect high-quality caption generation, this method utilizes multiple prompts to enable MLLMs like LLaVA, ChatGPT, and Qwen to perceive complex remote sensing scenes from diverse expert perspectives. The approach employs DINOv3 for extracting dense visual representations of land covers and features a Dynamic MixExperts module to adaptively integrate the most effective textual semantics. Additionally, a Linguistic Query Guided Attention mechanism is constructed to leverage textual semantic information for guiding visual features, ensuring precise segmentation. The proposed method demonstrates superior performance across three public semantic segmentation remote sensing datasets, marking a significant advancement in multimodal fusion for geospatial analysis. This development highlights the growing integration of advanced AI models in interpreting complex environmental data, offering improved accuracy in scene understanding and segmentation tasks within the field of computer vision and remote sensing.
cs.AI updates on arXiv.org