Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Methods of Documentation VII: EMR01:30

Methods of Documentation VII: EMR

881
Electronic Medical Records (EMRs) primarily center around electronically documenting patients' health information within a single healthcare organization or practice. They contain essential clinical data related to a patient's medical history, diagnoses, medications, treatment plans, lab results, and other pertinent information relevant to the specific encounter or episode of care. EMRs are designed to streamline documentation and workflow processes within individual healthcare...
881
Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

164
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
164
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

461
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
461
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

74
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
74
Multiple Regression01:25

Multiple Regression

3.1K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.1K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

92
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
92

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Access to inpatient video-EEG monitoring for patients with frequent seizure-related emergency visits.

Epilepsy research·2026
Same author

Patient- and Physician-Identified Considerations for Clinical Implementation of a New Risk-Based Prediction Tool to Guide Surveillance Mammography in Breast Cancer Survivors.

JCO oncology practice·2026
Same author

Performance of Statistical and Machine Learning Risk Prediction Models for Advanced Breast Cancers.

Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive Oncology·2026
Same author

Racial disparities and utilization trends of first-line targeted therapies for metastatic breast cancer.

JNCI cancer spectrum·2026
Same author

Trends in prevalence and incidence of diabetic retinal disease in patients with type 1 and type 2 diabetes mellitus.

Journal of diabetes and its complications·2026
Same author

Missingness in Eligibility Criteria for Target Trial Emulation in EHR With Survival Outcomes.

Statistics in medicine·2026

Related Experiment Video

Updated: Aug 12, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.6K

Performance of Multiple Imputation Using Modern Machine Learning Methods in Electronic Health Records Data.

Kylie Getz1, Rebecca A Hubbard2,3, Kristin A Linn2

  • 1From the Department of Biostatistics and Epidemiology, School of Public Health, Rutgers University, Piscataway, NJ.

Epidemiology (Cambridge, Mass.)
|February 1, 2023
PubMed
Summary

Advanced machine learning imputation methods, like denoising autoencoders, do not outperform traditional techniques for electronic health records (EHR) data with missingness not at random. These complex methods may introduce bias and spurious precision in epidemiological studies.

More Related Videos

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.2K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Related Experiment Videos

Last Updated: Aug 12, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.6K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.2K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Area of Science:

  • Epidemiology
  • Biostatistics
  • Health Informatics

Background:

  • Missing data are prevalent in electronic health records (EHR)-derived datasets.
  • Missingness patterns in EHR data often correlate with healthcare utilization, leading to complex 'missing not at random' mechanisms.
  • Machine learning imputation methods have been proposed as superior alternatives to traditional approaches, even for 'missing not at random' data.

Purpose of the Study:

  • To compare the performance of multiple imputation using chained equations, random forests, and denoising autoencoders.
  • To evaluate bias and precision of hazard ratio estimates using plasmode simulations from a metastatic urothelial carcinoma EHR database.
  • To assess performance under various missing data proportions and mechanisms (missing completely at random, missing at random, missing not at random).

Main Methods:

  • Plasmode simulations utilizing a nationwide de-identified EHR database.
  • Comparison of three multiple imputation techniques: chained equations, random forests, and denoising autoencoders.
  • Analysis focused on bias and precision of hazard ratio estimates across different missing data scenarios.

Main Results:

  • Chained equations and random forests showed low bias and similar precision under 'missing completely at random' conditions.
  • Denoising autoencoders exhibited higher bias than chained equations and random forests under 'missing at random' conditions.
  • All methods, including denoising autoencoders, displayed substantial bias under 'missing not at random', which escalated with increased missing data.

Conclusions:

  • Denoising autoencoders offer no advantage for multiple imputation in EHR-based epidemiologic studies.
  • Denoising autoencoders may lead to overfitting and inadequate confounder control.
  • Advanced imputation methods do not resolve bias from 'missing not at random' data and can create a false sense of precision.