AI privacy improves by auditing vector retrieval process safety
Beyond Private Training: The New Landscape of AI Privacy
Cryptography and SecurityInformation Retrieval
Summary
When AI systems look up information stored as vectors, deleted items may still affect how results are found even if they don’t appear in outputs. The authors studied this hidden exposure risk and made a tool called TSD-AUDIT to check and fix it. Their tool ensures deleted vectors aren’t even considered during search, not just hidden afterward, which keeps deleted data truly private. This improves search accuracy after deleting parts of the data, showing that just filtering results is not enough to protect privacy.
What this means in practice
- •For machine learning engineers: Ensure deleted data cannot influence search steps in vector-based AI systems to improve privacy compliance and retrieval accuracy.
- •For search infrastructure teams: Audit and certify that approximate nearest neighbor indexes fully exclude deleted entries at query time to maintain data protection guarantees.
Authors
Sean Culatana, Kang Li
Abstract
Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib's mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3--42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.