TPAW: Team-Based Self-Play with Dual Adaptive Weighting for LLM Fine-Tuning
Researchers have introduced Team-based self-Play with dual Adaptive Weighting (TPAW), a novel algorithm designed to enhance the alignment of Large Language Models (LLMs) in fully self-supervised settings. Addressing critical limitations in current self-training methods, such as sensitivity to synthetic data quality and ineffective optimization due to diminishing response gaps, TPAW employs a team-based framework. In this system, the current policy model collaborates with and competes against historical checkpoints to ensure stable and efficient optimization. The method incorporates two adaptive weighting mechanisms: a response reweighting scheme to adjust target response importance and a player weighting strategy to dynamically modulate team member contributions. Initialized from a Supervised Fine-Tuning (SFT) model, TPAW iteratively refines alignment without requiring additional human supervision. Experimental results indicate that TPAW consistently outperforms existing baselines across various base models and LLM benchmarks. The research team has made the code publicly available to support further development and verification in the field of artificial intelligence.
Wire timeline
TPAW: Team-Based Self-Play with Dual Adaptive Weighting for LLM Fine-Tuning
Researchers have introduced Team-based self-Play with dual Adaptive Weighting (TPAW), a novel algorithm designed to enhance the alignment of Large Language Models (LLMs) in fully self-supervised settings. Addressing critical limitations in current self-training methods, such as sensitivity to synthetic data quality and ineffective optimization due to diminishing response gaps, TPAW employs a team-based framework. In this system, the current policy model collaborates with and competes against historical checkpoints to ensure stable and efficient optimization. The method incorporates two adaptive weighting mechanisms: a response reweighting scheme to adjust target response importance and a player weighting strategy to dynamically modulate team member contributions. Initialized from a Supervised Fine-Tuning (SFT) model, TPAW iteratively refines alignment without requiring additional human supervision. Experimental results indicate that TPAW consistently outperforms existing baselines across various base models and LLM benchmarks. The research team has made the code publicly available to support further development and verification in the field of artificial intelligence.
cs.AI updates on arXiv.org