Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
Researchers have proposed a new framework for optimizing data selection in the instruction tuning of large language models (LLMs). Addressing the limitations of static, task-agnostic weighting schemes, this method dynamically adapts data selection to specific downstream tasks and model capabilities. By leveraging in-context learning signals on compact validation sets, the framework identifies optimal multi-indicator weights without requiring full-scale fine-tuning, thus serving as an efficient performance proxy. Experimental results across major model families, including Mistral, Qwen, and Llama, demonstrate that the approach achieves performance comparable to or exceeding full-dataset tuning while utilizing only 30% of training samples on the GSM8K benchmark. The study also highlights a critical trade-off between semantic diversity and logical complexity in reasoning tasks, underscoring the necessity of joint task-model adaptation for efficient LLM development. This research contributes to reducing computational costs in AI training while maintaining high fidelity in model performance.
Wire timeline
Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
Researchers have proposed a new framework for optimizing data selection in the instruction tuning of large language models (LLMs). Addressing the limitations of static, task-agnostic weighting schemes, this method dynamically adapts data selection to specific downstream tasks and model capabilities. By leveraging in-context learning signals on compact validation sets, the framework identifies optimal multi-indicator weights without requiring full-scale fine-tuning, thus serving as an efficient performance proxy. Experimental results across major model families, including Mistral, Qwen, and Llama, demonstrate that the approach achieves performance comparable to or exceeding full-dataset tuning while utilizing only 30% of training samples on the GSM8K benchmark. The study also highlights a critical trade-off between semantic diversity and logical complexity in reasoning tasks, underscoring the necessity of joint task-model adaptation for efficient LLM development. This research contributes to reducing computational costs in AI training while maintaining high fidelity in model performance.
cs.AI updates on arXiv.org