LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
Researchers have introduced LiteMedCoT-VL, a novel pipeline designed to enhance the reasoning capabilities of compact vision-language models (VLMs) for medical applications. Addressing the deployment limitations of large models on portable clinical devices, this method transfers chain-of-thought reasoning from a massive 235B parameter teacher model to efficient 2B parameter student models using LoRA-based fine-tuning. Unlike existing knowledge distillation techniques that only transfer final answers, LiteMedCoT-VL incorporates explanation-enriched training data to enable multi-step reasoning essential for interpretable clinical decision support. The system operates without image captions, simulating real-world scenarios where physicians interpret medical images directly. On the PMC-VQA benchmark, the model achieved 64.9% accuracy, surpassing the zero-shot Qwen3-VL-4B baseline by 11.0 percentage points and outperforming all published baselines. Visual grounding analysis confirmed the model relies on actual image content rather than textual priors. This advancement demonstrates that small models with distilled reasoning can match or exceed larger counterparts, facilitating the integration of advanced AI into resource-constrained healthcare environments. The source code has been made publicly available to support further research and development in medical AI.
Wire timeline
LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
Researchers have introduced LiteMedCoT-VL, a novel pipeline designed to enhance the reasoning capabilities of compact vision-language models (VLMs) for medical applications. Addressing the deployment limitations of large models on portable clinical devices, this method transfers chain-of-thought reasoning from a massive 235B parameter teacher model to efficient 2B parameter student models using LoRA-based fine-tuning. Unlike existing knowledge distillation techniques that only transfer final answers, LiteMedCoT-VL incorporates explanation-enriched training data to enable multi-step reasoning essential for interpretable clinical decision support. The system operates without image captions, simulating real-world scenarios where physicians interpret medical images directly. On the PMC-VQA benchmark, the model achieved 64.9% accuracy, surpassing the zero-shot Qwen3-VL-4B baseline by 11.0 percentage points and outperforming all published baselines. Visual grounding analysis confirmed the model relies on actual image content rather than textual priors. This advancement demonstrates that small models with distilled reasoning can match or exceed larger counterparts, facilitating the integration of advanced AI into resource-constrained healthcare environments. The source code has been made publicly available to support further research and development in medical AI.
cs.AI updates on arXiv.org