Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Calibration Curves: Correlation Coefficient01:10

Calibration Curves: Correlation Coefficient

5.0K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
5.0K
Deindividuation00:57

Deindividuation

22.4K
Deindividuation is a form of social influence on an individual’s behavior such that the individual engages in unusual or non-normal behavior while in a group setting. Why? Because in these group settings, the individual no longer sees themselves as an individual anymore, disinhibiting their behavior and personal restraint.
22.4K
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)01:20

¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)

1.3K
When proton-coupled carbon-13 spectra are simplified by a broadband proton decoupling technique, structural information about the coupled protons is lost. Distortionless enhancement by polarization transfer (DEPT) is a technique that provides information on the number of hydrogens attached to each carbon in a molecule. While the DEPT experiment utilizes complex pulse sequences, the pulse delay and flip angle are specifically manipulated. The resulting signals have different phases depending on...
1.3K
Difference from Background: Limit of Detection01:05

Difference from Background: Limit of Detection

9.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
9.0K
Self-Discrepancy and Its Effects01:29

Self-Discrepancy and Its Effects

515
Self-discrepancy theory explains how people compare their actual self to their ideal and ought selves and how mismatches between these self-guides can lead to emotional distress. Developed by E. Tory Higgins, the theory distinguishes among three components of self-concept: the actual self, the ideal self, and the ought self. These refer respectively to how individuals perceive themselves, how they aspire to be, and how they believe they are obligated to be. Emotional well-being, self-esteem,...
515
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

7.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
7.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Probiotics in football (soccer): a survey on practitioner's current perceptions and practices.

Science & medicine in football·2026
Same author

NLP in Support of Pharmacovigilance: QUality Adverse Drug Reaction AcTIve Control (QUADRATIC).

Clinical pharmacology and therapeutics·2026
Same author

Detection of Antithrombotic-Related Bleeding in Older Inpatients: Multicenter Retrospective Study Using Structured and Unstructured Electronic Health Record Data.

Journal of medical Internet research·2026
Same author

Melatonin Improves Intestinal Barrier Impairment in a Mouse Model of Autism Spectrum Disorder.

Biology·2025
Same author

From literature to biodiversity data: mining arthropod organismal traits with machine learning.

Biodiversity data journal·2025
Same author

Safety First: A Comprehensive Review of Nutritional Supplements for Hair Loss in Breast Cancer Patients.

Nutrients·2025

Related Experiment Video

Updated: May 2, 2026

Quantifying Intermembrane Distances with Serial Image Dilations
07:45

Quantifying Intermembrane Distances with Serial Image Dilations

Published on: September 28, 2018

8.7K

Measuring the gap: correlating synthetic-to-real drift with PHI de-identification performance.

Joseph Cornelius1,2, Fabio Rinaldi3

  • 1Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI-SUPSI), Via la Santa 1, Lugano-Viganello, CH-6962, Ticino, Switzerland. joseph.cornelius@idsia.ch.

Genomics & Informatics
|May 1, 2026
PubMed
Summary

Synthetic clinical notes generated by large language models (LLMs) aid de-identification in low-resource settings, but their utility depends on data source and quality control. Drift estimation can improve synthetic data alignment.

Keywords:
Clinical text de-identificationDistributional driftLarge language modelsSynthetic clinical data

More Related Videos

Sample Drift Correction Following 4D Confocal Time-lapse Imaging
10:04

Sample Drift Correction Following 4D Confocal Time-lapse Imaging

Published on: April 12, 2014

15.7K
Measurement of the Directional Information Flow in fNIRS-Hyperscanning Data using the Partial Wavelet Transform Coherence Method
08:42

Measurement of the Directional Information Flow in fNIRS-Hyperscanning Data using the Partial Wavelet Transform Coherence Method

Published on: September 3, 2021

2.9K

Related Experiment Videos

Last Updated: May 2, 2026

Quantifying Intermembrane Distances with Serial Image Dilations
07:45

Quantifying Intermembrane Distances with Serial Image Dilations

Published on: September 28, 2018

8.7K
Sample Drift Correction Following 4D Confocal Time-lapse Imaging
10:04

Sample Drift Correction Following 4D Confocal Time-lapse Imaging

Published on: April 12, 2014

15.7K
Measurement of the Directional Information Flow in fNIRS-Hyperscanning Data using the Partial Wavelet Transform Coherence Method
08:42

Measurement of the Directional Information Flow in fNIRS-Hyperscanning Data using the Partial Wavelet Transform Coherence Method

Published on: September 3, 2021

2.9K

Area of Science:

  • Natural Language Processing
  • Clinical Informatics
  • Machine Learning

Background:

  • Electronic health records (EHRs) require de-identification for patient privacy.
  • Publicly available de-identification training data are scarce and often lack consistent documentation styles.
  • Large language models (LLMs) offer a potential solution for generating synthetic clinical notes.

Purpose of the Study:

  • To evaluate the impact of lexical and semantic drift on protected health information (PHI) tagger performance.
  • To assess the utility of LLM-generated synthetic clinical notes for training de-identification models.
  • To determine if synthetic data reflects real-world clinical note distributions.

Main Methods:

  • Generated synthetic clinical notes using five generator LLMs and one judge LLM.
  • Fine-tuned de-identification models on real, synthetic, and mixed corpora.
  • Evaluated model performance on three external benchmarks using a harmonized label schema.
  • Assessed correlation between drift measures and out-of-distribution F1 score.

Main Results:

  • Models trained on broad, clinically relevant data sources outperformed those trained on legal or narrowly synthetic data.
  • Synthetic data, despite lacking some real-world distributional properties, proved useful in low-resource scenarios.
  • Compact distributional and embedding-based drift measures showed moderate correlation with out-of-distribution F1 score.

Conclusions:

  • LLM-generated synthetic clinical notes can augment scarce real-world data for de-identification tasks.
  • Careful selection of training data sources and quality control through drift estimation are crucial for maximizing synthetic data utility.
  • Drift estimation offers a practical method for improving synthetic data quality and alignment in clinical text de-identification.