Related Experiment Video
Updated: Feb 9, 2026

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
Published on: May 16, 2022
Differentially private data augmentation via LLM generation with discriminative and distribution-aligned filtering
Yiping Song1, Juhua Zhang2, Zhiliang Tian2
1College of Science, National University of Defense Technology, No.109, Deya Road, Kaifu District, Changsha, Hunan, 410073, China.
Abstract:
Data augmentation (DA) is a widely adopted approach for mitigating data insufficiency. Conducting DA in private domains requires privacy-preserving text generation, including anonymization or perturbation applied to sensitive textual data. The above methods lack formal protection guarantees. Existing Differential Privacy (DP) learning methods provide theoretical guarantees by adding calibrated noise to models or outputs. However, the large output space and model scales in text generation require substantial noise, which severely degrades synthesis quality. In this paper, we transfer DP-based synthetic sample generation to DP-based sample discrimination. Specifically, we propose a DP-based DA framework with a large language model (LLM) and a DP-based discriminator for private-domain text generation. Our key idea is to (1) leverage LLMs to generate large-scale high-quality samples, (2) select synthesized samples fitting the private domain, and (3) align the label distribution with the private domain. To achieve this, we use knowledge distillation to construct a DP-based discriminator: teacher models, accessing private data, guide a student model to select samples under calibrated noise. A DP-based tutor further constrains the label distribution of synthesized samples with a low privacy budget. We theoretically analyze the privacy guarantees and empirically validate our method on three medical text classification datasets, showing that our DP-synthesized samples significantly outperform state-of-the-art DP fine-tuning baselines in utility.
Related Concept Videos
Data: Types and Distribution
Distributions in...
Stereotypes, Prejudice, and Discrimination
Passive Filters
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Active Filters
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...

