Related Experiment Video
Updated: Jun 25, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A refined approach for evaluating small datasets via binary classification using machine learning
Steffen Steinert1,2, Verena Ruf1, David Dzsotjan1
1Chair of Physics Education, Ludwig-Maximilians-Universität München (LMU Munich), Munich, Germany.
Machine learning analysis of small datasets in education research requires careful evaluation. This study introduces a refined approach using permutation tests and nested cross-validation to ensure reliable, unbiased results for binary classification tasks.
Area of Science:
- Machine Learning
- Statistical Analysis
- Educational Research
Background:
- Classical statistical methods are often complemented or replaced by machine learning (ML).
- Small datasets, common in fields like education research, pose challenges related to bias and spurious findings.
- Evaluating ML performance on limited data requires specialized techniques to ensure reliability.
Purpose of the Study:
- To present a refined methodology for evaluating binary classification performance using ML on small datasets.
- To address issues of bias and chance in ML model evaluation within data-limited research contexts.
- To provide guidelines for selecting appropriate evaluation metrics for small-dataset ML applications.
Main Methods:
- Implementation of a non-parametric permutation test to assess the generalizability of ML model results.
- Utilization of repeated nested cross-validation for bias-free and reliable performance estimation.
- Comparative analysis of various evaluation metrics, including the Matthews correlation coefficient.
Main Results:
- Repeated nested cross-validation demonstrates minimal bias and high reliability, with results largely independent of chance.
- The permutation test effectively quantifies the probability of results generalizing to new, unseen data.
- The Matthews correlation coefficient is identified as a robust metric for binary classification when classes have equal importance, showing low bias and chance of coincidental success.
Conclusions:
- A combination of evaluation metrics is recommended for training and assessing ML classifiers to leverage their respective strengths.
- The proposed approach, incorporating permutation tests and nested cross-validation, is crucial for accurate ML analysis of small datasets.
- Avoiding biases is paramount when applying machine learning techniques to small datasets, particularly in sensitive research areas like education.
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...

