Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the vertical...
What Are Outliers?01:12

What Are Outliers?

Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Lipidomic profile of meningiomas harboring different NF2 mutation status.

Metabolomics : Official journal of the Metabolomic Society·2026
Same author

Microextraction Technologies as Exposomic Sensors for On-Site Environmental Air Monitoring of Volatile Organic Compounds: A Review of Commercially Available Technologies.

Molecules (Basel, Switzerland)·2026
Same author

Insights into diet, psychological distress, and personality traits among patients with lower-extremity lymphedema and overweight/obesity in comparison to patients with lifestyle-induced overweight/obesity and patients with normal body weight.

Obesity research & clinical practice·2025
Same author

Spatial variability of pollution source contributions during two (2012-2013 and 2018-2019) sampling campaigns at ten sites in Los Angeles basin.

Environmental pollution (Barking, Essex : 1987)·2024
Same author

Association between skin lymphangiogenesis parameters and arterial hypertension status in patients: An observational study.

Advances in clinical and experimental medicine : official organ Wroclaw Medical University·2024
Same author

Quantification and Detection of Ground Garlic Adulteration Using Fourier-Transform Near-Infrared Reflectance Spectra.

Foods (Basel, Switzerland)·2023

Related Experiment Video

Updated: Jul 16, 2026

Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

How to construct a multiple regression model for data with missing elements and outlying objects.

Ivana Stanimirova1, Sven Serneels, Pierre J Van Espen

  • 1Department of Chemometrics, The University of Silesia, Katowice, Poland.

Analytica Chimica Acta
|March 28, 2007
PubMed
Summary

Robust regression techniques using expectation maximization effectively model data with missing values and outliers. This approach offers a reliable method for building accurate regression models, outperforming standard methods in challenging datasets.

More Related Videos

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
10:46

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data

Published on: December 9, 2015

Related Experiment Videos

Last Updated: Jul 16, 2026

Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
10:46

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data

Published on: December 9, 2015

Area of Science:

  • Statistics
  • Data Science
  • Machine Learning

Background:

  • Traditional regression models struggle with datasets containing missing values and outliers.
  • Robust statistical methods are essential for reliable data analysis when data quality is compromised.

Purpose of the Study:

  • To demonstrate the effectiveness of robust multiple regression techniques within the expectation maximization framework.
  • To compare the performance of partial least squares and partial robust M-regression models using expectation maximization.

Main Methods:

  • Implementation of robust multiple regression techniques using the expectation maximization algorithm.
  • Comparative analysis involving partial least squares and partial robust M-regression.
  • Validation on simulated datasets with varying percentages of missing data and outliers, and on a real-world dataset.

Main Results:

  • The expectation maximization framework successfully accommodates missing data and outliers in regression modeling.
  • Partial robust M-regression within the expectation maximization framework showed strong performance.
  • The proposed methods yielded satisfactory regression models, evaluated by trimmed root mean squared errors.

Conclusions:

  • Robust multiple regression techniques integrated with expectation maximization provide a powerful solution for modeling incomplete and outlier-prone data.
  • The demonstrated methodology enhances the reliability and accuracy of regression models in practical applications.
  • This approach is valuable for data scientists and statisticians dealing with imperfect datasets.