LPT: Less-overfitting Prompt Tuning for Vision-Language Models
Researchers have introduced LPT, a novel framework designed to address the critical issue of overfitting in vision-language models (VLMs) during prompt tuning. While prompt learning is an efficient method for transferring VLMs to downstream tasks, it often suffers from severe overfitting, which significantly degrades generalization capabilities. The proposed LPT framework mitigates this by employing CLIP to filter out fine-grained foreground information that contributes to overfitting, thereby guiding prompts with basic visual concepts. Additionally, the method incorporates a Structural Preservation (SP) constraint at the feature level to align the model's feature space structure with frozen CLIP, ensuring plasticity and effective reshaping during optimization. Complementing this, a Hierarchical Logit (HL) constraint is applied at the output layer to manage overall class information. Extensive experiments across various benchmarks, including base-to-novel, cross-dataset transfer, and domain generalization scenarios, demonstrate that LPT significantly improves generalization and effectively alleviates overfitting compared to current state-of-the-art methods, marking a substantial advancement in efficient VLM adaptation.
Wire timeline
LPT: Less-overfitting Prompt Tuning for Vision-Language Models
Researchers have introduced LPT, a novel framework designed to address the critical issue of overfitting in vision-language models (VLMs) during prompt tuning. While prompt learning is an efficient method for transferring VLMs to downstream tasks, it often suffers from severe overfitting, which significantly degrades generalization capabilities. The proposed LPT framework mitigates this by employing CLIP to filter out fine-grained foreground information that contributes to overfitting, thereby guiding prompts with basic visual concepts. Additionally, the method incorporates a Structural Preservation (SP) constraint at the feature level to align the model's feature space structure with frozen CLIP, ensuring plasticity and effective reshaping during optimization. Complementing this, a Hierarchical Logit (HL) constraint is applied at the output layer to manage overall class information. Extensive experiments across various benchmarks, including base-to-novel, cross-dataset transfer, and domain generalization scenarios, demonstrate that LPT significantly improves generalization and effectively alleviates overfitting compared to current state-of-the-art methods, marking a substantial advancement in efficient VLM adaptation.
cs.AI updates on arXiv.org