CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
Researchers Zhipeng Liu and Chunbo Luo have introduced CrossVL, a novel framework designed to enhance vision-language models (VLMs) for cross-view object detection. Standard VLMs often struggle with the geometric and complexity differences between ground-level and aerial viewpoints, leading to significant performance degradation. CrossVL addresses this by combining Complexity-Aware Pathway Aggregation (CPA), which routes visual features through specific pathways based on scene complexity, and Paired Curriculum Learning (PCL), which utilizes semantic consistency in synchronized image pairs for stable training supervision. Tested on the MAVREC dataset, CrossVL improved the aerial mean Average Precision (mAP) of the Florence-2 model from 58.66% to 61.03%. Furthermore, it reduced the performance gap between ground and aerial views from 8.63 percentage points to 6.65 percentage points and achieved a 3.3x reduction in variance across random seeds. This study highlights the importance of coordinated architectural and training adaptations for robust cross-view detection in computer vision applications.
Wire timeline
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
Researchers Zhipeng Liu and Chunbo Luo have introduced CrossVL, a novel framework designed to enhance vision-language models (VLMs) for cross-view object detection. Standard VLMs often struggle with the geometric and complexity differences between ground-level and aerial viewpoints, leading to significant performance degradation. CrossVL addresses this by combining Complexity-Aware Pathway Aggregation (CPA), which routes visual features through specific pathways based on scene complexity, and Paired Curriculum Learning (PCL), which utilizes semantic consistency in synchronized image pairs for stable training supervision. Tested on the MAVREC dataset, CrossVL improved the aerial mean Average Precision (mAP) of the Florence-2 model from 58.66% to 61.03%. Furthermore, it reduced the performance gap between ground and aerial views from 8.63 percentage points to 6.65 percentage points and achieved a 3.3x reduction in variance across random seeds. This study highlights the importance of coordinated architectural and training adaptations for robust cross-view detection in computer vision applications.
cs.AI updates on arXiv.org