Related Experiment Video
Updated: Sep 15, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Q performance 2: Toward Unbiased Selection of Machine-Learning Regression Models
Arkaprava Banerjee1, Kunal Roy1
1Drug Theoretics and Cheminformatics Laboratory (DTC Lab), Department of Pharmaceutical Technology, Jadavpur University, Kolkata700032, India.
Abstract:
Selecting a single best machine-learning regression model from a set of competing models can be challenging. While models selected based on cross-validation performance do not guarantee good predictions on external data, models selected solely on external validation performance do not ascertain precise predictions for other external sets. Therefore, we propose three quantitative metrics to guide modelers in selecting the best model using the modeling set information only. Three quantitative data sets of varying sizes and complexities were considered. Each data set was randomly split into a modeling set and an independent test set. The modeling set was further split thrice to generate training and validation sets. Various machine-learning models were developed and validated against the validation sets. Our proposed metrics were computed for each model using only training and validation set performances. The novel metric values guided the selection of the best models, which also demonstrated expected performance on the independent test set. Furthermore, successful applications of our framework on two additional benchmark data sets demonstrated its wider generalizability.
Related Concept Videos
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Randomized Experiments
Simple randomization
Simple...
Goodness-of-Fit Test
