Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
A new research paper published on arXiv introduces the concept of "polylogue" to analyze how large language models (LLMs) process reasoning. The study builds on prior findings that LLMs encode behavioral traits, or "personas," as linear directions in activation space. Unlike previous approaches that treated these persona vectors as static controls, this research monitors them as dynamic signals during generation. The authors define polylogue as the time series of alignments between persona vectors and hidden activations. Experiments across four open-weight models demonstrate that polylogue features can predict correctness on the MMLU-Pro benchmark competitively with low-dimensional activation baselines while remaining interpretable. Furthermore, the study identifies specific latent directions for modulation at different response stages. By implementing a simple paragraph-conditioned intervention, the researchers improved accuracy on three of the four tested models. This work positions polylogue as a valuable tool for monitoring and intervening in LLM reasoning processes, suggesting that stage-aware latent steering is a promising avenue for enhancing model control and performance.
Wire timeline
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
A new research paper published on arXiv introduces the concept of "polylogue" to analyze how large language models (LLMs) process reasoning. The study builds on prior findings that LLMs encode behavioral traits, or "personas," as linear directions in activation space. Unlike previous approaches that treated these persona vectors as static controls, this research monitors them as dynamic signals during generation. The authors define polylogue as the time series of alignments between persona vectors and hidden activations. Experiments across four open-weight models demonstrate that polylogue features can predict correctness on the MMLU-Pro benchmark competitively with low-dimensional activation baselines while remaining interpretable. Furthermore, the study identifies specific latent directions for modulation at different response stages. By implementing a simple paragraph-conditioned intervention, the researchers improved accuracy on three of the four tested models. This work positions polylogue as a valuable tool for monitoring and intervening in LLM reasoning processes, suggesting that stage-aware latent steering is a promising avenue for enhancing model control and performance.
cs.AI updates on arXiv.org