DeepTumorVQA: A Hierarchical 3D CT Benchmark for Medical VLMs
Researchers have introduced DeepTumorVQA, a new hierarchical benchmark designed to evaluate medical vision-language models (VLMs) and AI agents in tumor diagnosis. Unlike existing benchmarks that rely on single accuracy scores, DeepTumorVQA decomposes 3D CT reasoning into four distinct stages: recognition, measurement, visual reasoning, and medical reasoning. The dataset comprises 476,000 questions across 42 clinical subtypes derived from 9,262 3D CT volumes. It supports both direct reasoning and tool-augmented environments, allowing models to utilize external segmentation and measurement tools. Evaluations of over 30 model configurations revealed that reliable quantitative measurement is the primary bottleneck for VLMs, though tool augmentation significantly mitigates this issue. The study highlights that leveraging medical knowledge and tools introduces new challenges but provides ground-truth step-by-step traces to supervise agents and reduce failures. This work offers a concrete roadmap for future medical AI development by clarifying where and why models fail in complex diagnostic tasks. All data and code have been made publicly available to facilitate further research in medical computer vision and artificial intelligence.
Wire timeline
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Medical VLMs
Researchers have introduced DeepTumorVQA, a new hierarchical benchmark designed to evaluate medical vision-language models (VLMs) and AI agents in tumor diagnosis. Unlike existing benchmarks that rely on single accuracy scores, DeepTumorVQA decomposes 3D CT reasoning into four distinct stages: recognition, measurement, visual reasoning, and medical reasoning. The dataset comprises 476,000 questions across 42 clinical subtypes derived from 9,262 3D CT volumes. It supports both direct reasoning and tool-augmented environments, allowing models to utilize external segmentation and measurement tools. Evaluations of over 30 model configurations revealed that reliable quantitative measurement is the primary bottleneck for VLMs, though tool augmentation significantly mitigates this issue. The study highlights that leveraging medical knowledge and tools introduces new challenges but provides ground-truth step-by-step traces to supervise agents and reduce failures. This work offers a concrete roadmap for future medical AI development by clarifying where and why models fail in complex diagnostic tasks. All data and code have been made publicly available to facilitate further research in medical computer vision and artificial intelligence.
cs.AI updates on arXiv.org