Related Experiment Video
Updated: Jan 9, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.9K
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation.
Summary
This study introduces a machine learning framework combining feature selection and data augmentation for accurate and interpretable omics-based classification, even with limited patient samples.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning in Healthcare
Background:
- Omics datasets present challenges in classification due to high dimensionality and limited samples.
- Current models often lack interpretability, hindering trust and reproducibility in clinical applications.
Purpose of the Study:
- To develop a machine learning framework integrating feature selection and data augmentation for improved omics-based classification.
- To enhance model transparency and reliability in disease classification using omics data.
Main Methods:
- A machine learning pipeline combining feature selection with data augmentation techniques.
- Bootstrap analysis on the E-MTAB-8026 dataset across six binary classification scenarios.
- Evaluation of cross-validated performance and generalization to larger test sets.
Main Results:
- The proposed framework achieves high classification accuracy and improved interpretability.
- Cross-validated performance on small datasets was maintained when applied to larger test sets.
- Synthetic data augmentation positively impacted model generalization, especially with limited sample availability.
Conclusions:
- The framework offers a balance between accuracy and feature selection for reliable omics-based classification.
- Data augmentation is crucial for enhancing generalization in omics studies with scarce data.
- This approach supports the development of explainable, reproducible diagnostic tools for clinical decision-making.
More Related Videos
Related Concept Videos
Synthetic Biology
5.5K
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
5.5K
Classification of Systems-I
540
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
540
Classification of Systems-II
446
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
446
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Aggregates Classification
953
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
953

