Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

13.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
13.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.4K
3.4K
Reliability and Validity01:29

Reliability and Validity

13.6K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.6K
Confidence Coefficient01:24

Confidence Coefficient

10.1K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
10.1K
Randomized Experiments01:13

Randomized Experiments

8.7K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
8.7K
Data Validation01:03

Data Validation

6.2K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
6.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Accuracy of Visual Inspection Alone to Assess Joint Effusions of the Hand: Cross-Sectional Study.

Journal of medical Internet research·2026
Same author

The Alberta Quality Assessment Tool: Risk of Bias (AQAT:RoB) for the Evaluation of Medical Large Language Model Question-Answer Studies: Development and Pilot Validation.

Journal of medical Internet research·2026
Same author

The subtleties of abolishing "race correction" in clinical artificial intelligence.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

The Utility and Implications of Ambient Scribes in Primary Care.

JMIR AI·2024
Same author

Genetics Navigator: protocol for a mixed methods randomized controlled trial evaluating a digital platform to deliver genomic services in Canadian pediatric and adult populations.

BMJ open·2024
Same author

Empirical data drift detection experiments on real-world medical imaging data.

Nature communications·2024

Related Experiment Video

Updated: Dec 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

895

Exploring the Privacy-Preserving Properties of Word Embeddings: Algorithmic Validation Study.

Mohamed Abdalla1,2,3, Moustafa Abdalla4,5,6, Graeme Hirst1,2

  • 1Department of Computer Science, University of Toronto, Toronto, ON, Canada.

Journal of Medical Internet Research
|July 17, 2020
PubMed
Summary

Clinical word embeddings trained on deidentified data can still reveal patient information. Removing personal health information (PHI) before training is insufficient to guarantee privacy, necessitating alternative anonymization methods.

Keywords:
data anonymizationnatural language processingpersonal health recordsprivacy

More Related Videos

Decoding Natural Behavior from Neuroethological Embedding
08:00

Decoding Natural Behavior from Neuroethological Embedding

Published on: October 3, 2025

432
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.1K

Related Experiment Videos

Last Updated: Dec 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

895
Decoding Natural Behavior from Neuroethological Embedding
08:00

Decoding Natural Behavior from Neuroethological Embedding

Published on: October 3, 2025

432
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.1K

Area of Science:

  • Natural Language Processing
  • Medical Informatics
  • Data Privacy

Background:

  • Word embeddings are crucial for neural networks in representing language.
  • Publicly available clinical word embeddings were previously nonexistent.
  • This study is the first to investigate the privacy risks associated with releasing clinical word embedding models.

Purpose of the Study:

  • To demonstrate that traditional word embeddings from deidentified clinical text can expose sensitive patient data.
  • To explore the privacy implications of different word embedding methods applied to clinical corpora.

Main Methods:

  • Utilized word embeddings generated from 400,000 doctor-written consultation notes.
  • Experimented with three common word embedding techniques.
  • Assessed the privacy-preserving capabilities of each method.

Main Results:

  • Reconstructed up to 68.5% of patient names from embeddings trained on deidentified data.
  • Linked sensitive information to specific patients within the corpus.
  • Found that vector distances between patient names and billing codes reveal diagnostic information.

Conclusions:

  • Current methods of deidentifying clinical text for word embeddings pose significant patient privacy risks.
  • Personal health information (PHI) removal alone is an imperfect anonymization strategy.
  • Anonymization by PHI replacement may offer a more robust privacy-preserving alternative for clinical word embeddings.