Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
Researchers Tao Hu and Da-Wei Zhou have introduced DRAPE, a novel prompt-learning framework designed to enhance Multimodal Continual Instruction Tuning (MCIT) for Large Language Models. While existing methods rely on task-level module composition, they often fail to address the substantial variations in visual scenes and reasoning demands within individual tasks. DRAPE addresses this by synthesizing continuous, instance-specific soft prompts through cross-attention between textual instructions and visual patch features. This approach allows for dynamic adaptation to specific query-image pairs rather than relying on fixed prompt pools. To prevent catastrophic forgetting during sequential updates, the framework employs null-space gradient projection and CLIP-based prototype routing for efficient, task-label-free generator selection. Extensive experiments demonstrate that DRAPE achieves state-of-the-art performance compared to representative prompt-based and LoRA-based continual learning baselines. This advancement significantly improves the ability of Multimodal Large Language Models to acquire new capabilities in real-world deployment scenarios while maintaining stability across sequential tasks.
Wire timeline
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
Researchers Tao Hu and Da-Wei Zhou have introduced DRAPE, a novel prompt-learning framework designed to enhance Multimodal Continual Instruction Tuning (MCIT) for Large Language Models. While existing methods rely on task-level module composition, they often fail to address the substantial variations in visual scenes and reasoning demands within individual tasks. DRAPE addresses this by synthesizing continuous, instance-specific soft prompts through cross-attention between textual instructions and visual patch features. This approach allows for dynamic adaptation to specific query-image pairs rather than relying on fixed prompt pools. To prevent catastrophic forgetting during sequential updates, the framework employs null-space gradient projection and CLIP-based prototype routing for efficient, task-label-free generator selection. Extensive experiments demonstrate that DRAPE achieves state-of-the-art performance compared to representative prompt-based and LoRA-based continual learning baselines. This advancement significantly improves the ability of Multimodal Large Language Models to acquire new capabilities in real-world deployment scenarios while maintaining stability across sequential tasks.
cs.AI updates on arXiv.org