DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
Researchers have introduced Diagonal-Anchored Repulsive Knowledge Distillation (DARK), a novel framework designed to compress vision-language models for efficient on-device deployment in clinical settings. Addressing the performance degradation typical in knowledge distillation when significant capacity gaps exist between teacher and student models, DARK decomposes distillation loss into diagonal and off-diagonal terms. This approach anchors matched-pair alignment while transitioning the student model from imitating to repelling the teacher's non-target similarity structure. The method was demonstrated by distilling FetalCLIP, a 427M-parameter fetal ultrasound model, into MobileFetalCLIP, a compact 75M-parameter student model. MobileFetalCLIP features a visual encoder 26 times smaller than its teacher and runs in 1.6 milliseconds on an iPhone 16 Pro. Despite extreme compression, the student model matches or exceeds the teacher's performance on zero-shot benchmarks, including HC18 biometry validity and brain sub-plane F1 scores. Analysis indicates that DARK induces structured decorrelation, allowing the student to preserve confidence while avoiding inherited inter-class confusion, proving that controlled repulsion is more efficient than strict imitation under extreme compression constraints.
Wire timeline
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
Researchers have introduced Diagonal-Anchored Repulsive Knowledge Distillation (DARK), a novel framework designed to compress vision-language models for efficient on-device deployment in clinical settings. Addressing the performance degradation typical in knowledge distillation when significant capacity gaps exist between teacher and student models, DARK decomposes distillation loss into diagonal and off-diagonal terms. This approach anchors matched-pair alignment while transitioning the student model from imitating to repelling the teacher's non-target similarity structure. The method was demonstrated by distilling FetalCLIP, a 427M-parameter fetal ultrasound model, into MobileFetalCLIP, a compact 75M-parameter student model. MobileFetalCLIP features a visual encoder 26 times smaller than its teacher and runs in 1.6 milliseconds on an iPhone 16 Pro. Despite extreme compression, the student model matches or exceeds the teacher's performance on zero-shot benchmarks, including HC18 biometry validity and brain sub-plane F1 scores. Analysis indicates that DARK induces structured decorrelation, allowing the student to preserve confidence while avoiding inherited inter-class confusion, proving that controlled repulsion is more efficient than strict imitation under extreme compression constraints.
cs.AI updates on arXiv.org