DiffKT3D: Any2Any 3D Diffusion Model for Radiotherapy Dose Prediction
Researchers have introduced DiffKT3D, a novel unified 3D diffusion framework designed to enhance voxel-wise dose prediction in radiotherapy planning. Addressing the generalization challenges of bespoke models, this approach leverages prior knowledge from pretrained video diffusion models trained on large-scale vision datasets. The framework features an Any2Any conditional paradigm that utilizes modality-specific embeddings, allowing flexible conditioning across various clinical inputs such as CT scans, anatomical structures, and beam settings without the computational overhead of cross-attention. Additionally, the team implemented a reinforcement learning post-training mechanism guided by a clinically informed Scorecard, tailored to align with institutional treatment preferences. In comparative evaluations, DiffKT3D achieved state-of-the-art performance, reducing voxel-level Mean Absolute Error from 2.07 to 1.93 against the previous GDP-HMM challenge winner. The model also demonstrated superior image quality and better alignment with clinical preferences. This study highlights the potential of transferring diffusion priors through modality-aware conditioning and clinically aligned reinforcement learning to provide robust, generalizable solutions for diverse radiotherapy scenarios.
Wire timeline
DiffKT3D: Any2Any 3D Diffusion Model for Radiotherapy Dose Prediction
Researchers have introduced DiffKT3D, a novel unified 3D diffusion framework designed to enhance voxel-wise dose prediction in radiotherapy planning. Addressing the generalization challenges of bespoke models, this approach leverages prior knowledge from pretrained video diffusion models trained on large-scale vision datasets. The framework features an Any2Any conditional paradigm that utilizes modality-specific embeddings, allowing flexible conditioning across various clinical inputs such as CT scans, anatomical structures, and beam settings without the computational overhead of cross-attention. Additionally, the team implemented a reinforcement learning post-training mechanism guided by a clinically informed Scorecard, tailored to align with institutional treatment preferences. In comparative evaluations, DiffKT3D achieved state-of-the-art performance, reducing voxel-level Mean Absolute Error from 2.07 to 1.93 against the previous GDP-HMM challenge winner. The model also demonstrated superior image quality and better alignment with clinical preferences. This study highlights the potential of transferring diffusion priors through modality-aware conditioning and clinically aligned reinforcement learning to provide robust, generalizable solutions for diverse radiotherapy scenarios.
cs.AI updates on arXiv.org