Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Cluster Sampling Method01:20

Cluster Sampling Method

12.4K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.4K
Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.8K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Machine learning-based prognosis and early death prediction in de novo stage IV breast cancer patients with bone metastasis: a SEER database and multicentre retrospective study.

Scientific reports·2026
Same author

A health ecological model study of subjective cognitive decline among hypertensive patients in rural Shanxi, China.

Scientific reports·2026
Same author

Effectiveness of the entire blood transfusion process with and without a clinical decision support system: a retrospective before-and-after study.

Frontiers in medicine·2026
Same author

Multi-trajectory patterns of ADL, cognition, and depression with fall risk: evidence from a longitudinal study in China.

BMC geriatrics·2026
Same author

Antibody-drug conjugates in breast cancer: from mechanism to revolutionizing clinical practice.

Molecular cancer·2026
Same author

Food environment and cardiovascular kidney metabolic diseases mortality.

NPJ science of food·2026

Related Experiment Video

Updated: Aug 29, 2025

VDJ-Seq: Deep Sequencing Analysis of Rearranged Immunoglobulin Heavy Chain Gene to Reveal Clonal Evolution Patterns of B Cell Lymphoma
15:07

VDJ-Seq: Deep Sequencing Analysis of Rearranged Immunoglobulin Heavy Chain Gene to Reveal Clonal Evolution Patterns of B Cell Lymphoma

Published on: December 28, 2015

26.8K

Predict DLBCL patients' recurrence within two years with Gaussian mixture model cluster oversampling and multi-kernel

Meng Xing1, Yanbo Zhang1, Hongmei Yu1

  • 1Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China; Shanxi Provincial Key Laboratory of Major Diseases Risk Assessment, Taiyuan, China.

Computer Methods and Programs in Biomedicine
|September 11, 2022
PubMed
Summary

A new model predicts diffuse large B-cell lymphoma (DLBCL) relapse within two years. The GMM-SENN-MKL method improves prediction accuracy by handling data imbalance and heterogeneity for better patient outcomes.

Keywords:
Class imbalanceDLBCLGaussian mixture model clustering oversamplingMultiple kernel learningRecurrence prediction

More Related Videos

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

183
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Related Experiment Videos

Last Updated: Aug 29, 2025

VDJ-Seq: Deep Sequencing Analysis of Rearranged Immunoglobulin Heavy Chain Gene to Reveal Clonal Evolution Patterns of B Cell Lymphoma
15:07

VDJ-Seq: Deep Sequencing Analysis of Rearranged Immunoglobulin Heavy Chain Gene to Reveal Clonal Evolution Patterns of B Cell Lymphoma

Published on: December 28, 2015

26.8K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

183
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Area of Science:

  • Hematology
  • Oncology
  • Machine Learning in Medicine

Background:

  • Diffuse large B-cell lymphoma (DLBCL) is a common non-Hodgkin's lymphoma in adults.
  • Early relapse (within two years) is associated with a poor prognosis.
  • Accurate prediction of early relapse is crucial for timely intervention and improved patient management.

Purpose of the Study:

  • To develop a predictive model for identifying diffuse large B-cell lymphoma (DLBCL) patients at high risk of relapse within two years.
  • To provide clinicians with a tool for individualized treatment strategies based on relapse risk.
  • To address challenges of data class imbalance and clinical heterogeneity in DLBCL datasets.

Main Methods:

  • A secondary-level class imbalance method using Gaussian mixture model (GMM) clustering and resampling (SMOTEENN) was employed.
  • A multi-kernel support vector machine (SVM) was utilized to integrate heterogeneous clinical data.
  • The combined approach (GMM-SENN-MKL) was developed to identify patients likely to relapse within two years.

Main Results:

  • The Inverse Weighted -GMM +SMOTEENN method demonstrated superior performance compared to standard methods.
  • This approach showed an 8.75% increase in Area Under the ROC Curve (AUC) and reduced ECE and Brier scores.
  • Multiple kernel learning (MKL) achieved the best discrimination and calibration, with maximized AUC, accuracy, recall, precision, and F1 scores.

Conclusions:

  • The developed inverse weighted -GMM+SMOTEENN+MKL (GMM-SENN-MKL) method effectively handles class imbalance and data heterogeneity in DLBCL.
  • This model shows significant potential for accurately predicting recurrence in DLBCL patients within two years.
  • The GMM-SENN-MKL method offers a valuable tool for clinical decision-making and personalized treatment planning.