Related Experiment Video
Updated: May 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Unsupervised utility evaluation of text anonymization methods via neural language models.
Benet Manzanares-Salor1, David Sánchez1, Pierre Lison2
1Department of Computer Engineering and Mathematics, CYBERCAT-Center for Cybersecurity Research of Catalonia, ComSCIAM-Center for Computational Science and Applied Mathematics, Universitat Rovira i Virgili, Av. Paisos Catalans 26, Tarragona, 43007, Spain.
This study introduces a novel unsupervised metric for evaluating text anonymization utility. It uses neural language models to assess data usefulness, outperforming traditional metrics without human annotation.
Area of Science:
- Natural Language Processing
- Information Security
- Data Science
Background:
- Text anonymization balances privacy and data utility.
- Current evaluation uses precision/recall with human annotations, which have drawbacks.
- Existing metrics are ill-suited for privacy tasks, assuming a single ground truth and ignoring term semantics.
Purpose of the Study:
- Introduce the first unsupervised utility metric for anonymized texts.
- Develop a complete evaluation framework for text anonymization.
- Evaluate various anonymization methods comprehensively.
Main Methods:
- Utilize neural language models to quantify utility loss.
- Develop an unsupervised metric to assess anonymized text utility.
- Combine the new metric with a privacy metric for a complete framework.
Main Results:
- The proposed metric accurately captures anonymized text utility.
- The metric is more sensitive to varying anonymization intensities than precision.
- Empirical experiments on document clustering validate the metric's effectiveness.
Conclusions:
- The unsupervised metric overcomes limitations of precision and recall.
- The proposed framework enables robust text anonymization evaluation without human input.
- This facilitates better anonymization methods and data utility preservation.