Related Experiment Video
Updated: Jan 14, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
SysML: adaptive recommendation system for heterogeneous biomedical data preprocessing and modeling workflows
Jinhui Zhao1,2, Xinjie Zhao1,2, Chunxia Zhao1,2
1State Key Laboratory of Medical Proteomics, Dalian Institute of Chemical Physics, Chinese Academy of Sciences, No. 457 Zhongshan Road, Shahekou District, Dalian, Liaoning 116023, P.R. China.
Abstract:
The rapid growth of high-dimensional omics datasets in biomedical research has created an urgent need for computational frameworks that are both robust and adaptable to diverse data complexities. Although a wide range of specialized tools and algorithms are available, researchers often rely on trial-and-error approaches to select suitable analytical workflows, compromising both efficiency and reproducibility. In this study, we systematically benchmarked hundreds of algorithms-preprocessing combinations across three common biomedical data challenges, including small sample sizes, missing values, and class imbalance. Our results show that tree-based models (e.g. Gradient Boosting Decision Tree, XGBoost, and Random Forest) consistently perform well in scenarios involving small-sample and missing-data, while partial least squares discriminant analysis (PLS-DA) is more effective in addressing imbalanced classes. Unsupervised cluster methods such as K-means and Density-Based Spatial Clustering of Applications with Noise (DBSCAN) remain robust under moderate missingness, but their performance declines when missingness exceeds 10%. To support data-driven decision-making, we developed SysML, a web-based platform that recommends data-adaptive workflows based on dataset-specific characteristics. Validated on multiple real-world biomedical datasets, SysML demonstrated improvements in both model performance and workflow efficiency. Our findings underscore that adaptive data preprocessing, rather than algorithm choice alone, is critical for achieving reliable and reproducible machine learning applications in biomedicine.
More Related Videos
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model Approaches for Pharmacokinetic Data: Physiological Models
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...