Cataract-LMM: A Large-Scale Multi-Source Benchmark for Deep Learning in Surgical Video Analysis
Researchers have introduced Cataract-LMM, a comprehensive dataset designed to advance deep learning applications in computer-assisted cataract surgery. Addressing the limitations of existing resources regarding diversity and annotation depth, this new benchmark comprises 3,000 phacoemulsification surgery videos collected from two distinct surgical centers. The dataset features multi-layer annotations, including temporal surgical phases, instance segmentation of instruments and anatomical structures, instrument-tissue interaction tracking, and quantitative skill scores based on ICO-OSCAR and GRASIS competency rubrics. The study demonstrates the dataset's utility by benchmarking deep learning models across four key tasks: workflow recognition, scene segmentation, interaction tracking, and automated skill assessment. Additionally, it establishes a domain-adaptation baseline to evaluate model generalizability across different surgical centers. This resource aims to facilitate the development of robust, generalizable multi-task models for surgical workflow analysis and competency-based training, significantly contributing to the field of medical computer vision and AI-driven surgical assistance.
Wire timeline
Cataract-LMM: A Large-Scale Multi-Source Benchmark for Deep Learning in Surgical Video Analysis
Researchers have introduced Cataract-LMM, a comprehensive dataset designed to advance deep learning applications in computer-assisted cataract surgery. Addressing the limitations of existing resources regarding diversity and annotation depth, this new benchmark comprises 3,000 phacoemulsification surgery videos collected from two distinct surgical centers. The dataset features multi-layer annotations, including temporal surgical phases, instance segmentation of instruments and anatomical structures, instrument-tissue interaction tracking, and quantitative skill scores based on ICO-OSCAR and GRASIS competency rubrics. The study demonstrates the dataset's utility by benchmarking deep learning models across four key tasks: workflow recognition, scene segmentation, interaction tracking, and automated skill assessment. Additionally, it establishes a domain-adaptation baseline to evaluate model generalizability across different surgical centers. This resource aims to facilitate the development of robust, generalizable multi-task models for surgical workflow analysis and competency-based training, significantly contributing to the field of medical computer vision and AI-driven surgical assistance.
cs.AI updates on arXiv.org