PrivBench: A Holistic and Modular Benchmarking Platform for Evaluating Text-to-Text Privatization

Computation and Language

Summary

The authors created PrivBench, a new platform to help evaluate how well different methods protect private information in text by changing it. They point out that checking how private these methods really are is complicated, so their platform tests multiple important aspects in a structured way. PrivBench is designed to be easy to update and fair by letting researchers compete with real-time results shown publicly. It is free and open for anyone to use online.

Authors

Stephen Meisenbacher, Andreea-Elena Bodea, Ahmet Bilal Akın, Alexandra Klymenko, Jana Diesner, Florian Matthes

Abstract

Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text privatization, where the goal is to transform a sensitive input text into a privatized output by ideally masking (in)directly identifiable or otherwise private information. The evaluation of text-to-text privatization, however, is not straightforward, and the extant literature has utilized a myriad of techniques and metrics to quantify the privacy-preserving capabilities of privatization methods. Seeking to unify the evaluation of text-to-text privatization, we introduce PrivBench, a holistic and modular benchmarking platform for researchers and practitioners working on text privatization. PrivBench is holistic in that it evaluates privatization on a series of defined desiderata, which are structured into modules. PrivBench is not only modular but also extensible, allowing for future updates and benchmark versions. PrivBench is user-centered and promotes competition via real-time evaluation and a live public leaderboard. The platform is free to use and openly accessible at https://privbench.com/.