Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
Researchers have introduced a novel knowledge distillation framework designed to transfer complex 3D spatial reasoning capabilities from large-scale vision-language models (VLMs) to lightweight alternatives. Addressing the high computational costs associated with models like LLaVA-3D, the team distilled knowledge from a 7-billion parameter teacher model into a 2.29-billion parameter student model. This approach achieves an 8.7x reduction in inference latency and a threefold decrease in model size while retaining 54-72% of the original performance. A key innovation is the introduction of "Hidden CoT," which utilizes learnable latent tokens as an internal scratchpad for reasoning without requiring explicit chain-of-thought data. The framework employs VGGT as a vision encoder and features a multi-task distillation pipeline with uncertainty-aware loss weighting. Experimental results on ScanNet and 3D-FRONT datasets demonstrate strong spatial understanding, with 68-72% accuracy on proximity and contact tasks. This development enables efficient 3D scene question answering on resource-constrained platforms, marking the first use of latent scratchpad reasoning in distilled 3D VLMs.
Wire timeline
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
Researchers have introduced a novel knowledge distillation framework designed to transfer complex 3D spatial reasoning capabilities from large-scale vision-language models (VLMs) to lightweight alternatives. Addressing the high computational costs associated with models like LLaVA-3D, the team distilled knowledge from a 7-billion parameter teacher model into a 2.29-billion parameter student model. This approach achieves an 8.7x reduction in inference latency and a threefold decrease in model size while retaining 54-72% of the original performance. A key innovation is the introduction of "Hidden CoT," which utilizes learnable latent tokens as an internal scratchpad for reasoning without requiring explicit chain-of-thought data. The framework employs VGGT as a vision encoder and features a multi-task distillation pipeline with uncertainty-aware loss weighting. Experimental results on ScanNet and 3D-FRONT datasets demonstrate strong spatial understanding, with 68-72% accuracy on proximity and contact tasks. This development enables efficient 3D scene question answering on resource-constrained platforms, marking the first use of latent scratchpad reasoning in distilled 3D VLMs.
cs.AI updates on arXiv.org