Related Experiment Video
Updated: Jul 25, 2025

Using a Split-belt Treadmill to Evaluate Generalization of Human Locomotor Adaptation
Published on: August 23, 2017
We need to talk about standard splits
1City University of New York.
Abstract:
It is standard practice in speech & language technology to rank systems according to performance on a test set held out for evaluation. However, few researchers apply statistical tests to determine whether differences in performance are likely to arise by chance, and few examine the stability of system ranking across multiple training-testing splits. We conduct replication and reproduction experiments with nine part-of-speech taggers published between 2000 and 2018, each of which reports state-of-the-art performance on a widely-used "standard split". We fail to reliably reproduce some rankings using randomly generated splits. We suggest that randomly generated splits should be used in system comparison.
Related Concept Videos
Interpreting ¹H NMR Signal Splitting: The (n + 1) Rule
¹H NMR: Complex Splitting
Splitting diagrams or splitting tree diagrams are routinely used to depict such complex couplings. While drawing splitting diagrams, the splitting with the larger coupling constant is usually applied...
¹H NMR Signal Multiplicity: Splitting Patterns
Voltage Dividers
Kirchhoff's voltage law implies that the sum of the voltages across the resistors in series equals the source voltage. This means that the...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Fixing Double-strand Breaks

