Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.5K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Cluster Sampling Method01:20

Cluster Sampling Method

11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

57
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
57
Mass Spectrometry: Complex Analysis01:21

Mass Spectrometry: Complex Analysis

796
Mass spectrometry is an important technique for the identification of pure compounds. However, it has some limitations for the analysis of complex mixtures, often due to excessive fragmentation making the spectrum too complicated to decipher. Mass spectrometry can be combined with suitable separation methods in sequence, forming hyphenated methods, which are useful in the analysis of complex mixtures.
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
796
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

527
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
527

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A multi-task domain-adapted model to predict chemotherapy response from mutations in recurrently altered cancer genes.

iScience·2025
Same author

A Joint LLM-KG System for Disease Q&A.

IEEE journal of biomedical and health informatics·2025
Same author

Evaluating Explanations From AI Algorithms for Clinical Decision-Making: A Social Science-Based Approach.

IEEE journal of biomedical and health informatics·2024
Same author

ASTER: A Method to Predict Clinically Relevant Synthetic Lethal Genetic Interactions.

IEEE journal of biomedical and health informatics·2024
Same author

ExpertNet: A Deep Learning Approach to Combined Risk Modeling and Subtyping in Intensive Care Units.

IEEE journal of biomedical and health informatics·2023
Same author

scMoMaT jointly performs single cell mosaic integration and multi-modal bio-marker detection.

Nature communications·2023

Related Experiment Video

Updated: Jul 11, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.0K

Avoiding inferior clusterings with misspecified Gaussian mixture models.

Siva Rajesh Kasa1, Vaibhav Rajan2

  • 1School of Computing, National University of Singapore, COM1, 13, Computing Dr, Singapore, 117417, Singapore. kasa@u.nus.edu.

Scientific Reports
|November 6, 2023
PubMed
Summary

This study introduces a new clustering algorithm, SIA, to address issues with Gaussian Mixture Models (GMMs) when data is not perfectly Gaussian. SIA effectively avoids inferior clustering solutions, improving data analysis accuracy.

More Related Videos

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
12:11

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry

Published on: April 8, 2020

8.2K
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.5K

Related Experiment Videos

Last Updated: Jul 11, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.0K
Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
12:11

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry

Published on: April 8, 2020

8.2K
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.5K

Area of Science:

  • Data Science
  • Statistical Modeling
  • Machine Learning

Background:

  • Clustering is vital for exploratory data analysis across scientific fields.
  • Gaussian Mixture Models (GMMs) are popular for clustering but can fail with non-Gaussian or noisy data.
  • Misspecified GMMs can lead to incorrect classifications and inferior clustering solutions.

Purpose of the Study:

  • To identify and characterize inferior clustering solutions arising from misspecified GMMs.
  • To develop a novel GMM-based clustering algorithm that avoids both inferior and spurious solutions.
  • To improve the robustness and interpretability of clustering when underlying data distributions are unknown.

Main Methods:

  • Theoretical analysis of component asymmetry in misspecified GMMs.
  • Introduction of a new penalty term to mitigate inferior and spurious solutions.
  • Development of a novel clustering algorithm, SIA, incorporating a new model selection criterion.

Main Results:

  • Characterization of inferior clustering solutions, distinct from spurious solutions.
  • Demonstration that the proposed penalty term effectively avoids inferior solutions.
  • Empirical evidence showing SIA outperforms existing GMM-based methods in misspecified scenarios.

Conclusions:

  • The proposed SIA algorithm offers a robust solution for clustering with potentially misspecified GMMs.
  • SIA enhances the reliability of clustering by avoiding problematic solution types.
  • This work contributes to more accurate data analysis in diverse scientific applications.