Tencent Unveils 'Think in Games' Framework for Strategic AI Training
Tencent researchers have introduced a novel training framework named 'Think in Games' (TiG), designed to enhance the strategic reasoning capabilities of artificial intelligence models. By leveraging real match data from the popular multiplayer game Honor of Kings, the framework combines supervised learning, reinforcement learning, and Group Relative Policy Optimization (GRPO). A key finding of the study indicates that smaller language models can outperform significantly larger counterparts under specific conditions. For instance, the Qwen3-14B model achieved a 90.9% accuracy rate in strategic decision-making after just 2,000 training steps, surpassing the performance of Deepseek-R1, which recorded 86.7%. The researchers emphasize that TiG not only improves gameplay proficiency but also fosters explainable reasoning processes. While the initial application focuses on gaming environments, the team suggests that this approach holds substantial potential for broader applications in AI systems requiring complex strategic planning and logical deduction. This development highlights a shift towards more efficient model training methods that prioritize reasoning quality over sheer model size.
Wire timeline
Tencent Unveils 'Think in Games' Framework for Strategic AI Training
Tencent researchers have introduced a novel training framework named 'Think in Games' (TiG), designed to enhance the strategic reasoning capabilities of artificial intelligence models. By leveraging real match data from the popular multiplayer game Honor of Kings, the framework combines supervised learning, reinforcement learning, and Group Relative Policy Optimization (GRPO). A key finding of the study indicates that smaller language models can outperform significantly larger counterparts under specific conditions. For instance, the Qwen3-14B model achieved a 90.9% accuracy rate in strategic decision-making after just 2,000 training steps, surpassing the performance of Deepseek-R1, which recorded 86.7%. The researchers emphasize that TiG not only improves gameplay proficiency but also fosters explainable reasoning processes. While the initial application focuses on gaming environments, the team suggests that this approach holds substantial potential for broader applications in AI systems requiring complex strategic planning and logical deduction. This development highlights a shift towards more efficient model training methods that prioritize reasoning quality over sheer model size.
TechNode