Emergent Semantic Role Understanding in Language Models
A new research paper published on arXiv investigates how linguistic structure, specifically semantic role understanding, emerges within large language models. The study addresses a central question in artificial intelligence: whether the ability to identify 'who did what to whom' arises solely from pre-training or requires task-specific fine-tuning. By freezing decoder-only transformer models and training linear probes to extract semantic roles, the authors analyzed performance across various model scales. The findings indicate that frozen pre-trained representations contain substantial semantic role information, suggesting that this capability partially emerges from language modeling objectives alone. However, the performance of these frozen models does not fully match that of fine-tuned counterparts, indicating that emergence is incomplete without adaptation. Furthermore, the research reveals that as model scale increases, the internal implementation of semantic role structure shifts toward more distributed representations. This work provides critical insights into the interpretability of language models and the extent of supervision required for complex meaning representation.
Wire timeline
Emergent Semantic Role Understanding in Language Models
A new research paper published on arXiv investigates how linguistic structure, specifically semantic role understanding, emerges within large language models. The study addresses a central question in artificial intelligence: whether the ability to identify 'who did what to whom' arises solely from pre-training or requires task-specific fine-tuning. By freezing decoder-only transformer models and training linear probes to extract semantic roles, the authors analyzed performance across various model scales. The findings indicate that frozen pre-trained representations contain substantial semantic role information, suggesting that this capability partially emerges from language modeling objectives alone. However, the performance of these frozen models does not fully match that of fine-tuned counterparts, indicating that emergence is incomplete without adaptation. Furthermore, the research reveals that as model scale increases, the internal implementation of semantic role structure shifts toward more distributed representations. This work provides critical insights into the interpretability of language models and the extent of supervision required for complex meaning representation.
cs.AI updates on arXiv.org