Back to Research papers
Research paper index

The Privacy-Hallucination Tradeoff in Differentially Private Language Models

Krithika Ramesh, Krishna Pillutla, Danish Pruthi, Anjalie Field

arXiv:2609.00492Published August 31, 20260 citations
  • cs.AI
  • cs.CL

Abstract

Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tuned with DP tend to produce more hallucinations than non-DP counterparts, with increased severity as the privacy budget grows stricter. Second, we investigate model properties driving this tradeoff, demonstrating that DP mechanisms flatten output distributions, potentially redistributing probability mass toward factually incorrect alternatives. Third, through experiments where we control fact frequency in training data, we characterize how information frequency can reduce hallucination risks in DP models. Overall, our findings underscore the need for more nuanced privacy-preserving interventions that offer rigorous privacy guarantees without compromising factual accuracy.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.