Related Experiment Video
Updated: May 26, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A mixture model with a reference-based automatic selection of components for disease classification from protein
Ivica Kopriva1, Marko Filipović
1Division of Laser and Atomic R&D, Ruđer Bošković Institute, Bijenička cesta 54, 10000 Zagreb, Croatia. ikopriva@irb.hr
This study introduces a new bioinformatics method for analyzing biological samples using an additive mixture model. The approach effectively extracts disease-specific components for accurate cancer prediction, improving upon existing matrix factorization techniques.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Bioinformatics data analysis commonly employs linear mixture models to represent samples as additive mixtures of components.
- Blind matrix factorization methods can extract these components from mixture samples.
- Automatic selection of components for classification remains a challenge.
Purpose of the Study:
- To develop a novel additive mixture model for feature extraction in bioinformatics.
- To enable automatic selection of components for disease prediction without using label information.
- To improve the accuracy and interpretability of component-based analysis in cancer datasets.
Main Methods:
- Proposed a sample-by-sample additive mixture model utilizing sparseness-constrained factorization.
- Decomposed each sample into control-specific, case-specific, and neutral components.
- Determined the number of components via cross-validation and assigned features based on sample-specific thresholds.
- Enabled automatic component selection for classification without relying on prior label information.
Main Results:
- Applied the method to ovarian, prostate, and colon cancer datasets for disease prediction.
- Achieved high average sensitivities (e.g., 96.2% for ovarian cancer) and specificities (e.g., 99% for prostate cancer).
- Demonstrated improved prediction accuracy through reduced component complexity and automatic feature allocation.
Conclusions:
- The proposed sample-by-sample factorization method offers an alternative to simultaneous dataset factorization.
- Automatic component selection allows for direct use in classification, unlike standard methods.
- Disease-specific components can be interpreted as sub-modes for biomarker identification.
- The method enhances prediction accuracy by refining component selection and reducing complexity.
Related Concept Videos
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.

