Related Experiment Video
Updated: Feb 9, 2026

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
Published on: May 16, 2022
Differentially private data augmentation via LLM generation with discriminative and distribution-aligned filtering
Yiping Song1, Juhua Zhang2, Zhiliang Tian2
1College of Science, National University of Defense Technology, No.109, Deya Road, Kaifu District, Changsha, Hunan, 410073, China.
This study introduces a privacy-preserving data augmentation framework using large language models and differential privacy. It enhances text generation quality for private domains while maintaining formal privacy guarantees.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Data augmentation (DA) is crucial for mitigating data insufficiency but poses privacy risks in sensitive domains.
- Existing privacy-preserving text generation methods lack formal guarantees.
- Differential Privacy (DP) offers theoretical guarantees but often degrades synthesis quality in text generation.
Purpose of the Study:
- To develop a novel DP-based data augmentation framework for private-domain text generation.
- To improve the utility of synthetic text data while ensuring formal privacy protection.
- To address the limitations of existing DP methods in large-scale text generation.
Main Methods:
- Leveraging large language models (LLMs) for generating high-quality synthetic samples.
- Employing a DP-based discriminator, constructed via knowledge distillation, to select domain-fitting samples.
- Utilizing a DP-based tutor to align label distributions with the private domain under a low privacy budget.
Main Results:
- The proposed DP-based DA framework effectively generates high-quality, privacy-preserving text data.
- DP-synthesized samples significantly outperform state-of-the-art DP fine-tuning baselines in utility.
- Empirical validation on three medical text classification datasets demonstrates superior performance.
Conclusions:
- The novel DP-based DA framework offers a robust solution for privacy-preserving text generation in sensitive domains.
- This approach successfully balances data utility and formal privacy guarantees, outperforming existing methods.
- The method shows significant promise for applications requiring secure and effective data augmentation.
Related Concept Videos
Data: Types and Distribution
Distributions in...
Stereotypes, Prejudice, and Discrimination
Passive Filters
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Active Filters
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...

