AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
Researchers have introduced AnomalyClaw, a novel training-free agent designed to enhance Visual Anomaly Detection (VAD) across diverse domains such as industrial inspection and medical imaging. Addressing the limitations of single-domain models and unreliable single-inference judgments by Vision-Language Models (VLMs), AnomalyClaw employs a multi-round refutation process. The agent proposes candidate anomalies and verifies them against normal-sample references using a library of thirteen specialized tools for visual verification and expert probing. Tested on the CrossDomainVAD-12 benchmark, the system demonstrated significant performance improvements over direct inference methods, achieving macro-AUROC gains of up to 7.93 percentage points on various VLMs including GPT-5.5 and Qwen3.5-VL-27B. Additionally, an optional verbalized self-evolution extension allows the agent to build an online rulebook from internal disagreements without requiring oracle labels, further boosting accuracy. This approach highlights how agentic refutation can significantly improve the reasoning capabilities of VLMs in complex visual tasks, offering a robust solution for cross-domain anomaly detection challenges.
Wire timeline
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
Researchers have introduced AnomalyClaw, a novel training-free agent designed to enhance Visual Anomaly Detection (VAD) across diverse domains such as industrial inspection and medical imaging. Addressing the limitations of single-domain models and unreliable single-inference judgments by Vision-Language Models (VLMs), AnomalyClaw employs a multi-round refutation process. The agent proposes candidate anomalies and verifies them against normal-sample references using a library of thirteen specialized tools for visual verification and expert probing. Tested on the CrossDomainVAD-12 benchmark, the system demonstrated significant performance improvements over direct inference methods, achieving macro-AUROC gains of up to 7.93 percentage points on various VLMs including GPT-5.5 and Qwen3.5-VL-27B. Additionally, an optional verbalized self-evolution extension allows the agent to build an online rulebook from internal disagreements without requiring oracle labels, further boosting accuracy. This approach highlights how agentic refutation can significantly improve the reasoning capabilities of VLMs in complex visual tasks, offering a robust solution for cross-domain anomaly detection challenges.
cs.AI updates on arXiv.org