Related Experiment Video
Updated: Aug 8, 2026

Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
Benchmarking supervised classifiers within a design-of-experiments framework: Robust statistical inference in text
1The BioRobotics Institute, Scuola Superiore Sant'Anna, Pisa, Italy.
Benchmarking literary text classification requires rigorous methods to avoid misleading results. This study introduces a reproducible workflow using shared train/test splits and leakage-free preprocessing for reliable classifier evaluation.
Area of Science:
- Computational Linguistics
- Digital Humanities
- Statistical Learning
Background:
- Supervised classification with Bag-of-Words models is common in literary text mining.
- Existing benchmarking practices often suffer from methodological fragility, leading to unreliable results.
- Issues include improper feature selection, non-shared resampling splits, and over-reliance on average accuracy.
Purpose of the Study:
- To present a reproducible workflow for benchmarking supervised classifiers in text analysis.
- To apply a design-of-experiments framework for statistically coherent evaluation.
- To demonstrate the importance of rigorous methodology using Dante's Divina Commedia.
Main Methods:
- Implementation of a workflow combining leakage-free preprocessing and shared Monte Carlo train/test splits.
- Utilizing a design-of-experiments framework in R for text analysis.
- Employing paired statistical comparison, power assessment, and validation with mispaired designs.
Main Results:
- The proposed workflow ensures statistically coherent benchmarking of text classifiers.
- Leakage-free preprocessing is crucial, even for small biases, to prevent altered classifier rankings.
- Shared train/test splits are integral to the inferential design, not merely a technical detail.
Conclusions:
- Established text classification methods can yield reliable results under statistically sound conditions.
- Rigorous benchmarking is essential to avoid overstating performance differences between classifiers.
- Performance interpretation should consider transparency, stability, and relevance beyond accuracy alone.
Related Concept Videos
Group Design
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Experimental Designs
Comparing the Survival Analysis of Two or More Groups
Methods of Medium Optimization
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
