Related Experiment Video
Updated: May 8, 2026

Activity-based Training on a Treadmill with Spinal Cord Injured Wistar Rats
Published on: January 16, 2019
Reassessing the Role of Machine Learning in Clinical Prediction: A Benchmark of Predicting Walking Function after
Julia Bugajska1, Louis P Lukas1,2, Rüdiger Rupp3
1Department of Health Sciences and Technology (D-HEST), ETH Zurich, Zürich, Switzerland.
None:
In light of growing biomedical data, machine learning (ML) models offer tremendous potential for personalized prediction in medicine. However, the additional value provided by these computational tools should always be critically evaluated. Using the example of predicting walking ability after spinal cord injury (SCI), we highlight a popular scenario in which data-driven predictions are feasible but not clinically meaningful, as the task can be performed equally well by humans. We asked 11 human observers from diverse backgrounds (five researchers without clinical training but proven knowledge of SCI and the International Standards for Neurological Classification of SCI [ISNCSCI], and six neurologists experienced in SCI) to predict walking ability following SCI based on acute phase neurological status assessed by the ISNCSCI motor and sensory scores (≤40 days after injury [DAI]). Following an established clinical prediction rule, walking ability was defined by a binary label derived from the indoor walking ability subitem of the Spinal Cord Independence Measure. We compared the performance of human observers with extreme gradient boosting and logistic regression-based models, which represent popular approaches in clinical literature on SCI. Using 794 patients from the European Multicenter Study about SCI, we show that all approaches provide similar, excellent performance at population level (area under the receiver operating characteristic 0.93-0.95; accuracy 0.88-0.90). Importantly, predictions combined from multiple neurologists (accuracy: 0.89) were comparable with model-based predictions (accuracy: 0.88-0.90), whereas individual neurologists (accuracy: 0.79 [0.01]; mean [standard deviation]) were marginally outperformed by computational approaches (accuracy: 0.88-0.90), particularly for more heterogeneous incomplete injuries. Individual SCI researchers performed equally well compared with neurologists (accuracy: 0.78 [0.02]). Our results show that prediction of walking function following SCI, if described through a binary label, does not benefit from ML, as ensembles of clinical experts and researchers each achieve performance similar to a range of ML models and an established clinical prediction rule. This highlights two key considerations in clinical applications of data-driven prediction models in SCI: first, the importance of carefully choosing clinical outcome measures to target in a prediction task to achieve a true benefit, and second, the necessity of benchmarking human performance on specific tasks to determine whether meaningful differences are present.
More Related Videos
06:31Automated Gait Analysis to Assess Functional Recovery in Rodents with Peripheral Nerve or Spinal Cord Contusion Injury
Published on: October 6, 2020
07:28Author Spotlight: Using the MouseWalker to Quantify Locomotor Dysfunction in a Mouse Model of Spinal Cord Injury
Published on: March 24, 2023