Related Experiment Video
Updated: Dec 29, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Missing data imputation via the expectation-maximization algorithm can improve principal component analysis aimed at
Linda Malan1, Cornelius M Smuts1, Jeannine Baumgartner2
1Centre of Excellence for Nutrition, North-West University, Potchefstroom, South Africa.
Imputing missing data using the expectation-maximization (EM) algorithm before principal component analysis (PCA) improves biomarker and dietary pattern analysis. This method enhances accuracy, especially with smaller sample sizes, by correcting biased eigenvalues caused by missing values.
Area of Science:
- Nutritional science
- Biostatistics
- Data science
Background:
- Principal Component Analysis (PCA) is widely used for analyzing complex datasets, including biomarker profiles and dietary patterns.
- A common limitation in PCA is the presence of missing data, which can lead to biased results.
- The practice of imputing missing data before PCA is not consistently applied.
Purpose of the Study:
- To evaluate the effectiveness of the expectation-maximization (EM) algorithm for imputing missing data prior to PCA.
- To determine if EM imputation improves the derivation of biomarker profiles and dietary patterns.
- To assess the impact of missing data on PCA eigenvalues in nutritional research.
Main Methods:
- Numerical simulations were conducted to mimic real-world nutritional data with varying amounts of missing values.
- The expectation-maximization (EM) algorithm was used for missing data imputation.
- Principal Component Analysis (PCA) was applied to both original datasets (without missing values) and datasets with imputed missing values.
- Real-world data from the US National Health and Nutrition Examination Survey (NHANES) were analyzed.
Main Results:
- PCA applied to data with missing values resulted in biased eigenvalues compared to complete datasets.
- The bias in eigenvalues increased with a higher percentage of missing data.
- Imputing missing data using the EM algorithm before PCA led to eigenvalues that closely matched those from the complete dataset.
- These findings were validated using real-world NHANES data.
Conclusions:
- The EM algorithm is a reliable and advantageous method for imputing missing data before applying PCA in nutritional research.
- EM imputation effectively corrects eigenvalue bias caused by missing data, leading to more accurate biomarker profiles and dietary patterns.
- This technique is particularly beneficial for analyses involving smaller sample sizes.
More Related Videos
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

