Back to Research papers
Research paper index

On the Recoverability of Private Information Unlearning in Large Language Models

Shicheng Hu, Runzhi Tian, Ziqiao Wang, Yongyi Mao

arXiv:2608.29943Published August 30, 20260 citations
  • cs.LG
  • cs.CL

Abstract

Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to remove such information, but it remains unclear whether existing methods truly erase it or merely hide it within the model. A key challenge is quantifying the persistence of sensitive data under a unified evaluation framework. To address this, we construct a synthetic dataset containing fake private information and propose a white-box auditing framework to systematically assess whether claimed-forgotten information is genuinely removed. Using this framework, we evaluate five existing unlearning methods and find that a simple "inverse greedy" decoding -- selecting the least likely token at each step -- can recover supposedly forgotten private information. Our results reveal that current unlearning approaches often fail to fully eliminate sensitive information, highlighting the need for more reliable methods to ensure privacy in deployed LLMs.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.