Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

4.4K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.4K
Outliers and Influential Points01:08

Outliers and Influential Points

6.6K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.6K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

9.7K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.7K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

1.4K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
1.4K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

7.2K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
7.2K
What Are Outliers?01:12

What Are Outliers?

5.5K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
5.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Prevalence of homologous recombination repair genes alterations in metastatic castration-resistant prostate cancer, a multicentric study.

The French journal of urology·2026
Same author

A histo-clinical score to predict evolution to radioactive iodine-refractory of the follicular cell-derived thyroid carcinoma (PREDIRAIR): a single-centre, prospective, cohort study.

EClinicalMedicine·2026
Same author

Validating the 2022 WHO classification in thyroid carcinomas: Prospective evidence for predicting iodine resistance.

Histopathology·2026
Same author

Insights From Mutational and Transcriptomic Profiles in Epithelial-myoepithelial Carcinoma.

The American journal of surgical pathology·2026
Same author

Migrations and Tuberculosis: A comparative study of Mycobacterium tuberculosis genomic population structure in Brazil and Mozambique to historical triangular slave trade knowledge to reconstruct the origins of tuberculosis infections caused by Lineage 1 in Brazil.

Tuberculosis (Edinburgh, Scotland)·2026
Same author

Benign Sweat-Gland Tubular Adenoma Harboring an MYBL1::NFIB Fusion Gene.

The American Journal of dermatopathology·2026

Related Experiment Video

Updated: Mar 15, 2026

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

2.2K

A Bregman-proximal point algorithm for robust non-negative matrix factorization with possible missing values and

Stéphane Chrétien1, Christophe Guyeux2, Bastien Conesa3

  • 1National Physical Laboratory, Hampton Road, Teddington, Middlesex, UK. stephane.chretien@npl.co.uk.

BMC Bioinformatics
|September 3, 2016
PubMed
Summary

This study enhances Non-Negative Matrix Factorization (NMF) for datasets with missing or outlier data. The new method effectively imputes missing values and identifies outliers, improving feature extraction accuracy.

Keywords:
Feature extractionGene expression analysisNon-negative matrix factorizationOutliers and missing data

More Related Videos

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
12:39

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types

Published on: December 10, 2012

11.7K
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
13:44

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns

Published on: August 30, 2013

43.8K

Related Experiment Videos

Last Updated: Mar 15, 2026

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

2.2K
A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
12:39

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types

Published on: December 10, 2012

11.7K
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
13:44

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns

Published on: August 30, 2013

43.8K

Area of Science:

  • Machine Learning
  • Data Science
  • Bioinformatics

Background:

  • Non-Negative Matrix Factorization (NMF) is crucial for feature extraction.
  • Existing NMF methods struggle with missing or outlier data.
  • This study addresses limitations in NMF for real-world datasets.

Purpose of the Study:

  • To extend NMF applicability to datasets with missing and/or corrupted data.
  • To develop a robust method for simultaneous feature reconstruction and outlier detection.
  • To improve the accuracy of NMF in the presence of data imperfections.

Main Methods:

  • A novel Bregman proximal approach preserving nonnegativity.
  • Integration with the Augmented Lagrangian method.
  • Utilizing a sparsity-promoting ℓ1 penalty for outlier detection.

Main Results:

  • Simultaneous reconstruction of features and detection of outliers.
  • Effective handling of missing data imputation.
  • Demonstrated robustness against corrupted data points.

Conclusions:

  • The proposed method enhances NMF for imperfect datasets.
  • Applicable to gene expression data analysis in bladder cancer.
  • Offers a powerful tool for feature extraction in bioinformatics.