Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

280
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
280
Documentation of Nursing Diagnosis01:10

Documentation of Nursing Diagnosis

1.3K
The nurse documents nursing diagnoses and enters them into the patient record. The identified patient's nursing diagnosis is either written out with a plan of care or entered into the electronic health record.
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
1.3K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

369
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
369
Confounding in Epidemiological Studies01:27

Confounding in Epidemiological Studies

170
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
170
Introduction to Epidemiology01:26

Introduction to Epidemiology

736
Epidemiology, known as the cornerstone of public health, involves studying the distribution and determinants of health-related events in defined populations and applying these insights to control health issues. This is essential for understanding how diseases spread, identifying populations at greater risk, and implementing measures to control or prevent outbreaks. Epidemiology addresses not only infectious diseases but also non-communicable conditions like cancer and cardiovascular disease,...
736
Systematic Error: Methodological and Sampling Errors01:15

Systematic Error: Methodological and Sampling Errors

1.5K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Predictive models of adverse outcomes of overweight or obesity: a nationwide cohort validation.

BMC endocrine disorders·2026
Same author

Traditional Chinese Medicines for Metabolic Dysfunction-Associated Steatotic Liver Disease: A Mixed-Method Systematic Review.

Journal of evidence-based medicine·2026
Same author

Cerebellar pathway diffusion MRI measures are linked to core autism symptoms in early adolescents aged 9 to 11 years.

Brain structure & function·2026
Same author

Plant-Derived Thylakoids Potentiate Copper-Mediated Multimodal Cell Death via Hypoxia Alleviation for Synergistic Antitumor Therapy.

Small (Weinheim an der Bergstrasse, Germany)·2026
Same author

DDX24 exacerbates inflammation-induced immunosuppression in oral squamous cell carcinoma progression through IL-17 signaling pathway.

Archives of oral biology·2026
Same author

Liquid-liquid phase separation-driven regulated cell death: from molecular mechanisms to therapeutic strategies.

Cell death discovery·2026

Related Experiment Video

Updated: Jul 7, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.1K

Impact of possible errors in natural language processing-derived data on downstream epidemiologic analysis.

Zhou Lan1,2, Alexander Turchin2,3

  • 1Center for Clinical Investigation, Brigham & Women's Hospital, Boston, MA 02115, United States.

JAMIA Open
|December 28, 2023
PubMed
Summary

Simulated natural language processing (NLP) errors had a modest impact on epidemiologic study results. Small changes in effect estimates were observed, but the interpretation of findings remained consistent, suggesting NLP errors are unlikely to significantly affect study outcomes.

Keywords:
Monte Carlo methodelectronic health recordepidemiologynatural language processingoutcomes research

More Related Videos

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

15.9K
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.5K

Related Experiment Videos

Last Updated: Jul 7, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.1K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

15.9K
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.5K

Area of Science:

  • Epidemiology
  • Health Informatics
  • Computational Linguistics

Background:

  • Natural Language Processing (NLP) is increasingly used to extract variables for epidemiologic studies.
  • Potential errors in NLP algorithms could theoretically impact study results and conclusions.
  • Quantifying the impact of these errors is crucial for validating NLP applications in health research.

Purpose of the Study:

  • To assess the impact of simulated natural language processing (NLP) errors on the outcomes of epidemiologic research.
  • To determine if NLP-derived variables and their associated errors affect the interpretation of study findings.

Main Methods:

  • Utilized data from three outcomes research studies employing NLP for predictor variable generation.
  • Applied Monte Carlo simulations to create datasets with varying degrees of simulated NLP errors.
  • Fit original regression models to simulated datasets and compared coefficient estimates and significance to original results.

Main Results:

  • Mean changes in effect estimates for NLP-derived variables ranged from -21.9% to 4.12%.
  • Significance of the primary predictor-outcome relationship was largely maintained across simulations.
  • Confounder variable estimates showed minimal changes (0.27% to 2.27%), with no shifts in interpretation direction.

Conclusions:

  • Simulated errors in NLP-derived variables had a modest impact on epidemiologic study results.
  • The direction and significance of associations were generally preserved, indicating robustness of findings.
  • NLP errors are unlikely to substantially alter the conclusions of studies utilizing NLP-generated data.