Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Data: Types and Distribution01:19

Data: Types and Distribution

1.8K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
1.8K
Stereotypes, Prejudice, and Discrimination02:55

Stereotypes, Prejudice, and Discrimination

95.4K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
95.4K
Passive Filters01:27

Passive Filters

1.0K
Passive filters are utilized to shape the frequency spectrum of signals across a diverse array of applications. These filters, using only passive elements like resistors (R), inductors (L), and capacitors (C), are capable of selectively allowing or blocking certain frequency ranges without the need for external power sources.
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
1.0K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models01:06

Model Approaches for Pharmacokinetic Data: Distributed Parameter Models

249
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
249
Active Filters01:25

Active Filters

1.3K
Active filters are electronic circuits that use operational amplifiers (op-amps), resistors, and capacitors to filter out unwanted frequency components from a signal. A first-order low-pass active filter is designed to pass signals with a frequency lower than a certain cutoff frequency and attenuate frequencies higher than that cutoff frequency. The transfer function for a first-order low-pass active filter is:
1.3K
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

1.4K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Unlocking efficient near-infrared circularly polarized phosphorescence reaching 800 nm in cyclometalated Pt(II) complexes.

Chemical communications (Cambridge, England)·2026
Same author

Concentration-Responsive Gd-DOTA Nanomicelles via Chain-Length Engineering for Transporter-Independent Hepatobiliary Magnetic Resonance Imaging.

Langmuir : the ACS journal of surfaces and colloids·2026
Same author

A capacitive-piezoelectric hybrid MEMS microphone with signal fusion for enhancing signal-to-noise ratio.

Microsystems & nanoengineering·2026
Same author

Research Progress, Safety Regulation and Application Prospects in Health Food Development of Red Yeast Rice-Derived Bioactive Compounds: A Critical Narrative Review.

Foods (Basel, Switzerland)·2026
Same author

Sustainable Conversion of Plastic and Corrugated Paper into Ni@C Integrated N/S Dual-Doped Carbon Membranes for Accelerated Polysulfide Redox in Lithium Polysulfide Batteries.

Langmuir : the ACS journal of surfaces and colloids·2026
Same author

Formulation Screening of Lipid Nanoparticles Enhances mRNA Delivery to Retina.

Molecular pharmaceutics·2026

Related Experiment Video

Updated: Feb 9, 2026

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

17.5K

Differentially private data augmentation via LLM generation with discriminative and distribution-aligned filtering.

Yiping Song1, Juhua Zhang2, Zhiliang Tian2

  • 1College of Science, National University of Defense Technology, No.109, Deya Road, Kaifu District, Changsha, Hunan, 410073, China.

Neural Networks : the Official Journal of the International Neural Network Society
|February 7, 2026
PubMed
Summary

This study introduces a privacy-preserving data augmentation framework using large language models and differential privacy. It enhances text generation quality for private domains while maintaining formal privacy guarantees.

Keywords:
Data augmentationDifferential privacyLarge language modelSynthetic text generation

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K
Generation of Aligned Functional Myocardial Tissue Through Microcontact Printing
11:09

Generation of Aligned Functional Myocardial Tissue Through Microcontact Printing

Published on: March 19, 2013

11.6K

Related Experiment Videos

Last Updated: Feb 9, 2026

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

17.5K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K
Generation of Aligned Functional Myocardial Tissue Through Microcontact Printing
11:09

Generation of Aligned Functional Myocardial Tissue Through Microcontact Printing

Published on: March 19, 2013

11.6K

Area of Science:

  • Computer Science
  • Artificial Intelligence
  • Machine Learning

Background:

  • Data augmentation (DA) is crucial for mitigating data insufficiency but poses privacy risks in sensitive domains.
  • Existing privacy-preserving text generation methods lack formal guarantees.
  • Differential Privacy (DP) offers theoretical guarantees but often degrades synthesis quality in text generation.

Purpose of the Study:

  • To develop a novel DP-based data augmentation framework for private-domain text generation.
  • To improve the utility of synthetic text data while ensuring formal privacy protection.
  • To address the limitations of existing DP methods in large-scale text generation.

Main Methods:

  • Leveraging large language models (LLMs) for generating high-quality synthetic samples.
  • Employing a DP-based discriminator, constructed via knowledge distillation, to select domain-fitting samples.
  • Utilizing a DP-based tutor to align label distributions with the private domain under a low privacy budget.

Main Results:

  • The proposed DP-based DA framework effectively generates high-quality, privacy-preserving text data.
  • DP-synthesized samples significantly outperform state-of-the-art DP fine-tuning baselines in utility.
  • Empirical validation on three medical text classification datasets demonstrates superior performance.

Conclusions:

  • The novel DP-based DA framework offers a robust solution for privacy-preserving text generation in sensitive domains.
  • This approach successfully balances data utility and formal privacy guarantees, outperforming existing methods.
  • The method shows significant promise for applications requiring secure and effective data augmentation.