Related Experiment Video
Updated: Jun 25, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A comprehensive benchmarking of machine learning algorithms and dimensionality reduction methods for drug sensitivity
Lea Eckhart1, Kerstin Lenhof1, Lisa-Marie Rolli1
1Center for Bioinformatics, Saarland Informatics Campus, Saarland University, 66123, Saarland, Germany.
This study benchmarks machine learning (ML) methods and dimension reduction (DR) techniques for predicting anti-cancer drug response in cell lines. Optimized DR strategies enhance complex ML models, but simpler models can still achieve superior performance with fewer features.
Area of Science:
- Computational biology
- Genomics
- Pharmacology
Background:
- Precision oncology relies on identifying molecular biomarkers for targeted cancer treatments.
- Large cancer cell line datasets are crucial for understanding the link between cellular features and drug response.
- High-dimensional data necessitates machine learning (ML) for analysis, but algorithm and feature selection remain challenging.
Purpose of the Study:
- To comprehensively benchmark machine learning methods and dimension reduction techniques for predicting drug response metrics.
- To compare ML models based on statistical performance, runtime, and interpretability.
- To provide strategies for model performance assessment and complexity trade-offs.
Main Methods:
- Benchmarking of random forests, neural networks, boosting trees, and elastic nets.
- Application of nine dimension reduction (DR) approaches to feature sets.
- Training and evaluation using the Genomics of Drug Sensitivity in Cancer cell line panel for 179 anti-cancer compounds.
Main Results:
- Complex ML models show improved performance with optimized DR strategies.
- Standard ML models can outperform complex models even with reduced feature sets.
- Performance, runtime, and interpretability were key comparison metrics.
Conclusions:
- Dimension reduction is critical for optimizing complex machine learning models in precision oncology.
- Simpler machine learning models offer a viable and potentially superior alternative, especially when computational resources or interpretability are prioritized.
- The study provides a framework for evaluating and selecting appropriate ML and DR strategies for drug response prediction.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.

