Distinguishing Capability Elicitation from Creation in LLM Post-Training
A new academic paper submitted to arXiv challenges the conventional distinction between supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery in large language model post-training. The authors, Yuhao Li and Shengchao Liu, argue that this binary view is too coarse. Instead, they propose distinguishing between capability elicitation, which reweights behaviors already within a model's reachable set, and capability creation, which expands the model's accessible support. Using a free-energy perspective, the study operationalizes this by defining accessible support as the set of behaviors a model can practically produce under finite budgets. The framework suggests that both SFT and RL primarily function as reweighting mechanisms of a pretrained reference distribution, driven by demonstration or reward signals respectively. True capability creation, according to the authors, occurs only when post-training methods expand the reachable behavioral space through search, interaction, tool use, or new information incorporation. This theoretical contribution aims to refine how researchers evaluate the impact of different post-training strategies on AI model development.
Wire timeline
Distinguishing Capability Elicitation from Creation in LLM Post-Training
A new academic paper submitted to arXiv challenges the conventional distinction between supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery in large language model post-training. The authors, Yuhao Li and Shengchao Liu, argue that this binary view is too coarse. Instead, they propose distinguishing between capability elicitation, which reweights behaviors already within a model's reachable set, and capability creation, which expands the model's accessible support. Using a free-energy perspective, the study operationalizes this by defining accessible support as the set of behaviors a model can practically produce under finite budgets. The framework suggests that both SFT and RL primarily function as reweighting mechanisms of a pretrained reference distribution, driven by demonstration or reward signals respectively. True capability creation, according to the authors, occurs only when post-training methods expand the reachable behavioral space through search, interaction, tool use, or new information incorporation. This theoretical contribution aims to refine how researchers evaluate the impact of different post-training strategies on AI model development.
cs.AI updates on arXiv.org