Statistical auditing method measures identity risk of released text
Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees
Cryptography and SecurityComputation and Language
Summary
Text documents released online can reveal who wrote them when attackers use powerful language AI and extra information to guess identities. The authors created a method called Conformal Privacy Auditing that gives a statistical guarantee on whether an identity guess is correct for each document. This method works with different kinds of AI models and helps figure out how risky it is to share certain texts. It shows when the chance of being identified goes up or down depending on how the text is prepared and what extra information attackers have.
What this means in practice
- •For data privacy teams: Provide a reliable statistical measure that quantifies identity exposure risk for each released text document in privacy audits.
- •For api security teams: Test and compare the risk of user identity leaks when sharing text through proprietary language model APIs using a unified auditing framework.
Authors
Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu
Abstract
Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finite-sample statistical guarantees, while training-time protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing(CPA), a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.