Related Experiment Video
Updated: May 11, 2026

09:35
A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
18.3K
Towards a Biological Evaluation Framework for Oversampling (BEFO) gene expression data.
Kevin Fee1, Suneil Jain2, Ross G Murphy3
1Queen's University Belfast School of Electronics, Electrical Engineering and Computer Science, 16A Malone Rd, Belfast, BT9 5BN, Ulster, Northern Ireland, UK.
Journal of Biomedical Informatics
|October 19, 2025
Summary
This study introduces a Biological Evaluation Framework for Oversampling (BEFO) to improve machine learning models in biomedical research. BEFO ensures synthetic data reflects biological patterns, enhancing model accuracy and trustworthiness for clinical applications.
Area of Science:
- Biomedical research
- Machine learning applications
- Data science
Background:
- Machine learning (ML) models are increasingly used in biomedical research for improved diagnostics and prognostics.
- Biomedical datasets often exhibit class imbalance, leading to biased ML models.
- Existing oversampling techniques lack biological validation for synthetic data, limiting clinical applicability.
Purpose of the Study:
- To introduce the Biological Evaluation Framework for Oversampling (BEFO) to ensure synthetic gene expression data accurately reflects biological patterns.
- To mitigate bias in ML models and enhance the trustworthiness of predictions in clinical settings.
- To establish a new standard for evaluating synthetic data in biomedical ML.
Main Methods:
- Developed a ranking method for synthetic samples based on Weighted Gene Co-expression Network Analysis (WGCNA) gene co-expression clusters.
- Constructed random forests to assess synthetic sample alignment with biological clusters.
- Included only synthetic samples demonstrating higher importance than real samples.
Main Results:
- The BEFO framework improved the biological feasibility of oversampled datasets by an average of 11%.
- Classification performance improved by an average of 9% compared to state-of-the-art methods.
- Evaluated across six real-world gene expression datasets using ten classification algorithms.
Conclusions:
- The proposed ML oversampling framework enhances biological relevance and predictive performance.
- BEFO offers a robust method for validating synthetic data in biomedical ML.
- This approach improves the reliability of ML decision support systems in clinical practice.

