Related Experiment Video
Updated: Aug 8, 2026

Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
Benchmarking supervised classifiers within a design-of-experiments framework: Robust statistical inference in text
1The BioRobotics Institute, Scuola Superiore Sant'Anna, Pisa, Italy.
Abstract:
Supervised classification based on Bag-of-Words representations is widely used in literary text mining, yet benchmarking practices often remain methodologically fragile. Common problems include feature selection before train/test separation, comparisons based on non-shared resampling splits, and inferential conclusions drawn from average accuracy alone. These choices may produce misleading significance and overstate marginal performance differences between classifiers. This article presents a reproducible workflow for benchmarking supervised classifiers within a design-of-experiments framework for text analysis in R, using Dante's Divina Commedia as an empirical testbed. The workflow combines leakage-free preprocessing, shared Monte Carlo train/test splits, paired statistical comparison, pilot-based power assessment, and validation through deliberately mispaired designs. Elastic-net multinomial logistic regression and linear support vector machine are used as competing classifiers. The objective is not to propose a new classifier, but to show how established methods can be applied under statistically coherent conditions. Shared train/test splits should be treated as part of the inferential design rather than as a technical detail of model fitting. Leakage-free preprocessing remains essential even when bias appears numerically small, because leakage can alter classifier ranking and distort inferential interpretation. Small performance differences should be interpreted alongside model transparency, inferential stability, and practical relevance, rather than through accuracy alone.
Related Concept Videos
Group Design
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Experimental Designs
Comparing the Survival Analysis of Two or More Groups
Methods of Medium Optimization
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
