Study Reveals Involuntary Information Leakage in Large Language Models
A new research paper titled 'Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing' highlights significant security vulnerabilities in frontier large language models (LLMs). The study investigates whether LLMs can effectively compartmentalize sensitive information, such as system prompts or secret data, when instructed to keep them hidden. Researchers tested five leading models by providing a secret word and instructing the model to write a story without revealing it. Although the secret word never appeared literally in the output, all models leaked the information thematically through topic choices, imagery, and settings. A secondary model successfully identified the secret in up to 79% of cases, a rate significantly higher than chance. The leakage was found to scale with model size and remain detectable even when models actively attempted to avoid the secret. However, this vulnerability disappeared in short-form writing like jokes. The findings suggest that attending to a secret opens an information channel that current LLMs cannot fully close, posing risks for deployments requiring strict data confidentiality and prompt protection.
Wire timeline
Study Reveals Involuntary Information Leakage in Large Language Models
A new research paper titled 'Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing' highlights significant security vulnerabilities in frontier large language models (LLMs). The study investigates whether LLMs can effectively compartmentalize sensitive information, such as system prompts or secret data, when instructed to keep them hidden. Researchers tested five leading models by providing a secret word and instructing the model to write a story without revealing it. Although the secret word never appeared literally in the output, all models leaked the information thematically through topic choices, imagery, and settings. A secondary model successfully identified the secret in up to 79% of cases, a rate significantly higher than chance. The leakage was found to scale with model size and remain detectable even when models actively attempted to avoid the secret. However, this vulnerability disappeared in short-form writing like jokes. The findings suggest that attending to a secret opens an information channel that current LLMs cannot fully close, posing risks for deployments requiring strict data confidentiality and prompt protection.
cs.AI updates on arXiv.org