Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Regression Analysis01:11

Regression Analysis

7.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.6K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.8K
Regression Toward the Mean01:52

Regression Toward the Mean

6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.4K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.4K
Survival Tree01:19

Survival Tree

337
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
337
Outliers and Influential Points01:08

Outliers and Influential Points

5.7K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Combining genome-wide polygenic scores with registry data for colorectal cancer risk-based screening.

British journal of cancer·2026
Same author

Glucocorticoid use among patients with rheumatoid arthritis during the first year of treatment-a cohort study.

Rheumatology international·2026
Same author

Real-world evidence on medication patterns in inflammatory bowel disease the first decade after diagnosis.

Journal of Crohn's & colitis·2026
Same author

Serial Aqueous Humor Proteomics in Diabetic Macular Edema: Correlation of Novel Targets with Retinal Fluid Volume.

Ophthalmology science·2026
Same author

Assessing the Accuracy of Symptoms and Adverse Events Reporting for Lung Cancer Treatment in the Danish National Patient Registry.

Clinical epidemiology·2026
Same author

The risk of multiple myeloma associated with daily low-dose exposure to ionising radiation from radon decay.

Cancer epidemiology·2026

Related Experiment Video

Updated: Dec 27, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.3K

Regression on imperfect class labels derived by unsupervised clustering.

Rasmus Froberg Brøndum, Thomas Yssing Michaelsen, Martin Bøgsted

    Briefings in Bioinformatics
    |March 4, 2020
    PubMed
    Summary

    This study addresses bias in regression analysis caused by misclassified class labels from unsupervised clustering. Methods like regression calibration reduce bias and improve confidence intervals for accurate effect parameter estimation.

    Keywords:
    Clusteringcancermachine learningstatisticssurvival analysis

    More Related Videos

    Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
    08:56

    Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

    Published on: January 13, 2023

    2.8K
    A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
    08:12

    A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

    Published on: March 1, 2022

    2.9K

    Related Experiment Videos

    Last Updated: Dec 27, 2025

    Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
    12:27

    Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

    Published on: February 15, 2017

    7.3K
    Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
    08:56

    Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

    Published on: January 13, 2023

    2.8K
    A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
    08:12

    A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

    Published on: March 1, 2022

    2.9K

    Area of Science:

    • Biostatistics
    • Machine Learning
    • Bioinformatics

    Background:

    • Unsupervised clustering is widely used to identify class labels for regression analysis.
    • Ignoring label misclassification introduces significant bias in estimated effect parameters.

    Purpose of the Study:

    • To develop and evaluate methods for correcting bias caused by misclassified class labels in regression analysis.
    • To improve the accuracy of effect parameter estimation in studies using unsupervised clustering.

    Main Methods:

    • Regression calibration
    • Misclassification simulation and extrapolation (MSX)
    • Application to Gaussian mixture models and gene expression data

    Main Results:

    • Both regression calibration and MSX methods significantly reduced bias.
    • Improved coverage of confidence intervals was observed when adjusting for misclassification.
    • Demonstrated effectiveness on simulated and real-world gene expression data.

    Conclusions:

    • Adjusting for class label misclassification is crucial for reliable regression analysis.
    • Regression calibration and MSX are effective strategies to mitigate bias in such applications.
    • The methods enhance the validity of findings in bioinformatics and other fields utilizing clustering.