autoPET3 Challenge Results: Automated Lesion Segmentation in Whole-Body PET/CT
This report details the design and outcomes of the third autoPET challenge, held at MICCAI 2024, focusing on automated lesion segmentation in whole-body PET/CT scans. The study addressed compositional generalization using a training dataset of 1,014 [18F]-FDG studies from University Hospital Tübingen and 597 PSMA studies from LMU University Hospital Munich, marking the largest public annotated PSMA PET/CT dataset. Seventeen teams submitted algorithms, primarily based on nnU-Net architectures. The top-performing algorithm achieved a mean Dice Similarity Coefficient of 0.66, significantly improving upon the baseline by reducing false-negative volumes. Key findings indicate that while in-domain multitracer segmentation is approaching human reader agreement, generalizing to unseen tracer-center combinations remains challenging due to systematic volume overestimation. The analysis highlights that data heterogeneity and case difficulty impact performance more than algorithmic choices among top contenders. This benchmark provides critical insights for advancing AI-driven medical imaging diagnostics.
Wire timeline
autoPET3 Challenge Results: Automated Lesion Segmentation in Whole-Body PET/CT
This report details the design and outcomes of the third autoPET challenge, held at MICCAI 2024, focusing on automated lesion segmentation in whole-body PET/CT scans. The study addressed compositional generalization using a training dataset of 1,014 [18F]-FDG studies from University Hospital Tübingen and 597 PSMA studies from LMU University Hospital Munich, marking the largest public annotated PSMA PET/CT dataset. Seventeen teams submitted algorithms, primarily based on nnU-Net architectures. The top-performing algorithm achieved a mean Dice Similarity Coefficient of 0.66, significantly improving upon the baseline by reducing false-negative volumes. Key findings indicate that while in-domain multitracer segmentation is approaching human reader agreement, generalizing to unseen tracer-center combinations remains challenging due to systematic volume overestimation. The analysis highlights that data heterogeneity and case difficulty impact performance more than algorithmic choices among top contenders. This benchmark provides critical insights for advancing AI-driven medical imaging diagnostics.
cs.AI updates on arXiv.org