Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

504
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
504
Regression Toward the Mean01:52

Regression Toward the Mean

7.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.3K
Random and Systematic Errors01:20

Random and Systematic Errors

16.3K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
16.3K
Random and Systematic Errors01:20

Random and Systematic Errors

977
977
Confounding in Epidemiological Studies01:27

Confounding in Epidemiological Studies

1.1K
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
1.1K
Errors occurring during blood pressure monitoring01:25

Errors occurring during blood pressure monitoring

1.9K
Blood pressure monitoring is a crucial clinical procedure in diagnosing and managing various cardiovascular conditions. Despite its significance, the accuracy of blood pressure measurements can be compromised by multiple factors, potentially leading to either falsely high or low readings. These inaccuracies are critical as they can significantly impact patient care. So, it is vital to understand these challenges deeply and adopt strategic approaches to minimize errors.
Several factors...
1.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Medication-Wide Association Study of Alzheimer's Disease and Related Dementias: Identifying Drug Candidates from Electronic Health Records through Explainable AI.

medRxiv : the preprint server for health sciences·2026
Same author

Characteristics and Outcomes of Over 1 Million Veterans With Heart Failure Phenotyped Using Artificial Intelligence Approaches: the National DCVA-HF Registry.

Journal of cardiac failure·2026
Same author

Target-Dose Versus Below-Target-Dose ACE Inhibitors and Lower Risk of Kidney Failure in U.S. Veterans with HFrEF.

European journal of heart failure·2026
Same author

Serum Magnesium and Outcomes in U.S. Veterans with Heart Failure.

The American journal of medicine·2026
Same author

Coding Fairness: Detecting Demographic-Related Coding Discrepancies in ICD Code Assignments.

AMIA ... Annual Symposium proceedings. AMIA Symposium·2026
Same author

Design of Personal Health Libraries for People Returning from Incarceration in the United States.

Proceedings of the ... Annual Hawaii International Conference on System Sciences. Annual Hawaii International Conference on System Sciences·2026

Related Experiment Videos

Beware the Little Foxes that Spoil the Vines: Small Inconsistencies in Clinical Data Can Distort Machine Learning

Abdolvahab Khademi1, Mark S Tuttle2, Qing Zeng-Treitler1

  • 1Biomedical Informatics Center, George Washington University, Washington, DC.

Fortune Journal of Health Sciences
|April 20, 2026
PubMed
Summary

Electronic Health Records (EHR) data noise significantly impairs predictive model accuracy and the ability to identify key health risk factors. Even small amounts of data inconsistency can obscure important signals, leading to misleading analysis results.

Keywords:
CodingData QualityInformation theoryInternational Classification of DiseasesNoise

Related Experiment Videos

Area of Science:

  • Health Informatics
  • Data Science
  • Biostatistics

Background:

  • Electronic Health Records (EHR) data are crucial for clinical research but often contain inconsistencies and inaccuracies.
  • The impact of data noise on predictive modeling and risk factor identification in EHR is frequently underestimated.

Purpose of the Study:

  • To investigate the effects of varying levels of random and non-random noise on the performance of common predictive models.
  • To assess how data noise influences the identification of significant predictors in EHR data.

Main Methods:

  • Simulated different levels of binary noise (random and non-random) in curated EHR data from the All of Us database.
  • Evaluated the performance of logistic regression, support vector machines, and gradient boosting models under varying noise conditions.

Main Results:

  • Increased data noise consistently reduced classification accuracy across all tested models.
  • Noise diminished the variance of variable impact scores, hindering the identification of key predictors, while means remained unchanged.
  • The observed effects were consistent across different models and noise types, indicating a data-driven issue.

Conclusions:

  • Even modest levels of noise in EHR data can obscure meaningful biological or clinical signals.
  • Standard performance metrics like accuracy and hazard ratios may be unreliable when analyzing noisy EHR data.
  • These findings have broad implications for the interpretation and reliability of analyses using real-world EHR data.