Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
A new research paper titled "Unlearners Can Lie" addresses critical safety concerns in large language model (LLM) unlearning. While unlearning aims to remove harmful data, existing methods often cause models to hallucinate or behave inconsistently, which researchers associate with dishonesty. The authors propose a formal definition of "unlearning honesty," requiring preserved utility on retained knowledge and consistent acknowledgment of limitations regarding forgotten data. They introduce a comprehensive evaluation suite covering utility, honesty, forgetting effectiveness, and refusal stability. An assessment of nine current methods across three mainstream families reveals that none meet these honesty standards. To address this, the team presents ReVa, a representation-alignment procedure that fine-tunes feature-randomized models. Experimental results show ReVa significantly outperforms existing techniques, achieving the highest rejection rate for forgotten knowledge questions and improving honesty on retained data. The study highlights the gap between current unlearning capabilities and safe deployment standards, offering a novel technical solution to enhance model trustworthiness. The authors have released their data and code to support further research in this emerging field of AI safety and model editing.
Wire timeline
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
A new research paper titled "Unlearners Can Lie" addresses critical safety concerns in large language model (LLM) unlearning. While unlearning aims to remove harmful data, existing methods often cause models to hallucinate or behave inconsistently, which researchers associate with dishonesty. The authors propose a formal definition of "unlearning honesty," requiring preserved utility on retained knowledge and consistent acknowledgment of limitations regarding forgotten data. They introduce a comprehensive evaluation suite covering utility, honesty, forgetting effectiveness, and refusal stability. An assessment of nine current methods across three mainstream families reveals that none meet these honesty standards. To address this, the team presents ReVa, a representation-alignment procedure that fine-tunes feature-randomized models. Experimental results show ReVa significantly outperforms existing techniques, achieving the highest rejection rate for forgotten knowledge questions and improving honesty on retained data. The study highlights the gap between current unlearning capabilities and safe deployment standards, offering a novel technical solution to enhance model trustworthiness. The authors have released their data and code to support further research in this emerging field of AI safety and model editing.
cs.AI updates on arXiv.org