Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
A new research paper submitted to arXiv investigates the extent to which large language models (LLMs) encode and retain authorial stylistic information within text embeddings, specifically focusing on the French language. The study addresses the growing capability of LLMs to convincingly imitate human writing styles, raising questions about the persistence of stylistic signals after AI-driven rewriting. Using a controlled literary dataset, the authors quantify the impact of stylistic variation by analyzing changes in embedding dispersion. The findings indicate that embeddings reliably capture distinct authorial features and that these signals persist even after the text has been rewritten by an LLM, although the process also introduces model-specific patterns. These analytical results provide significant insights for the development of methods to detect authorship imitation in the era of advanced language models. The research offers promising directions for distinguishing between original human writing and AI-generated or modified content, contributing to the broader field of computational linguistics and AI security.
Wire timeline
Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
A new research paper submitted to arXiv investigates the extent to which large language models (LLMs) encode and retain authorial stylistic information within text embeddings, specifically focusing on the French language. The study addresses the growing capability of LLMs to convincingly imitate human writing styles, raising questions about the persistence of stylistic signals after AI-driven rewriting. Using a controlled literary dataset, the authors quantify the impact of stylistic variation by analyzing changes in embedding dispersion. The findings indicate that embeddings reliably capture distinct authorial features and that these signals persist even after the text has been rewritten by an LLM, although the process also introduces model-specific patterns. These analytical results provide significant insights for the development of methods to detect authorship imitation in the era of advanced language models. The research offers promising directions for distinguishing between original human writing and AI-generated or modified content, contributing to the broader field of computational linguistics and AI security.
cs.AI updates on arXiv.org