MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
Researchers have introduced MMVIAD, the first continuous multi-view video dataset designed for industrial anomaly detection and understanding. Addressing the limitations of existing static image datasets, MMVIAD features object-centric 2-second inspection clips with significant camera motion, covering 48 object categories, 14 environments, and 6 structural anomaly types. The dataset supports multiple tasks, including anomaly detection, defect classification, object classification, and temporal localization. Evaluations indicate that current commercial and open-source video Multimodal Large Language Models (MLLMs) significantly underperform compared to human capabilities, particularly in fine-grained defect recognition. To bridge this gap, the authors developed a two-stage post-training pipeline comprising Perception-Structured Supervised Fine-Tuning (PS-SFT) and VISTA-GRPO. This process produced the VISTA model, which improved average scores across four tasks from 45.0 to 57.5 on unseen data, surpassing the performance of GPT-5.4. The source code for the dataset and model is publicly available, marking a significant advancement in automated manufacturing quality control and computer vision applications.
Wire timeline
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
Researchers have introduced MMVIAD, the first continuous multi-view video dataset designed for industrial anomaly detection and understanding. Addressing the limitations of existing static image datasets, MMVIAD features object-centric 2-second inspection clips with significant camera motion, covering 48 object categories, 14 environments, and 6 structural anomaly types. The dataset supports multiple tasks, including anomaly detection, defect classification, object classification, and temporal localization. Evaluations indicate that current commercial and open-source video Multimodal Large Language Models (MLLMs) significantly underperform compared to human capabilities, particularly in fine-grained defect recognition. To bridge this gap, the authors developed a two-stage post-training pipeline comprising Perception-Structured Supervised Fine-Tuning (PS-SFT) and VISTA-GRPO. This process produced the VISTA model, which improved average scores across four tasks from 45.0 to 57.5 on unseen data, surpassing the performance of GPT-5.4. The source code for the dataset and model is publicly available, marking a significant advancement in automated manufacturing quality control and computer vision applications.
cs.AI updates on arXiv.org