When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
A new research paper published on arXiv highlights significant privacy vulnerabilities in Large Language Models (LLMs) used for generating synthetic tabular data. The study, titled 'When Tables Leak,' demonstrates that both fine-tuned smaller models and prompted larger models often reproduce memorized numeric digit patterns from their training data, thereby compromising user privacy. To systematically assess this risk, the authors introduce LevAtt, a novel No-box Membership Inference Attack (MIA) that targets string sequences of numeric digits using only access to generated synthetic data. The attack revealed substantial privacy leakage across various models and datasets, occasionally achieving perfect membership classification on state-of-the-art systems. In response to these findings, the researchers propose two defensive methods, including a innovative sampling strategy that strategically perturbs digits during the generation process. Evaluations indicate that this approach effectively mitigates the identified attacks while maintaining high fidelity and utility of the synthetic data. This work underscores the urgent need for robust privacy defenses in LLM-based data generation frameworks to prevent unintended data leakage.
Wire timeline
When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
A new research paper published on arXiv highlights significant privacy vulnerabilities in Large Language Models (LLMs) used for generating synthetic tabular data. The study, titled 'When Tables Leak,' demonstrates that both fine-tuned smaller models and prompted larger models often reproduce memorized numeric digit patterns from their training data, thereby compromising user privacy. To systematically assess this risk, the authors introduce LevAtt, a novel No-box Membership Inference Attack (MIA) that targets string sequences of numeric digits using only access to generated synthetic data. The attack revealed substantial privacy leakage across various models and datasets, occasionally achieving perfect membership classification on state-of-the-art systems. In response to these findings, the researchers propose two defensive methods, including a innovative sampling strategy that strategically perturbs digits during the generation process. Evaluations indicate that this approach effectively mitigates the identified attacks while maintaining high fidelity and utility of the synthetic data. This work underscores the urgent need for robust privacy defenses in LLM-based data generation frameworks to prevent unintended data leakage.
cs.AI updates on arXiv.org