Related Experiment Video
Updated: Aug 21, 2026

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
External validation of a machine learning-based web application for personalized testing of objective functioning
Kenneth Arockia1, Massimo Bottini1, Anita M Klukowska2
1Machine Intelligence in Clinical Neuroscience & Microsurgical Neuroanatomy (MICN) Laboratory, Department of Neurosurgery, Clinical Neuroscience Center, University Hospital Zurich, University of Zurich, Frauenklinikstrasse 10, 8091, Zurich, Switzerland.
Purpose:
Simple generalized thresholds for assessing functional impairment in clinical testing are limited, as they fail to consider patient-specific properties such as age, body height, and body mass index. A previously developed machine learning-based model for personalized testing using the five-repetition sit-to-stand (5R-STS) test, that estimates personalized upper limits of normal (ULN) to identify objective functional impairment (OFI) was externally validated to evaluate its performance and generalizability across cohorts.
Methods:
Only healthy individuals were included in the study. After the machine learning-based model was applied to this external dataset, expected and observed 5R-STS test times were compared using standardized performance assessment metrics including root mean square error (RMSE), mean absolute error (MAE), and R2 values. Additionally, a Bland-Altman analysis was performed to assess agreement between observed and expected values. Validation of the expected ULN involved comparing the proportion of individuals exceeding their personalized thresholds with the corresponding proportion based on the generalized threshold. Subgroup analyses by test setting and country of residence were additionally performed, along with a graphical assessment of model performance.
Results:
Application of the model to 171 healthy individuals resulted in an RMSE of 2.33 (95% CI: 1.93 to 2.73) seconds, MAE of 1.70 (95% CI: 1.47 to 1.94) seconds, and R2 of 0.064 (95% CI: -0.25 to 0.15). The Bland-Altman analysis demonstrated a mean bias of -1.1 s. Based on the personalized ULNs, OFI was classified in 17.5% of individuals, compared to 6.4% when using the generalized threshold of 10.4 s. These analyses indicated limited external generalization, with acceptable approximation for some faster and mid-range test times but systematic underestimation of slower test times. Multivariable regression analyses showed that remote testing was independently associated with greater prediction bias and higher odds of personalized ULN exceedance compared with supervised testing (adjusted mean difference: - 0.90 s; adjusted OR: 10.7). Exploratory country-of-residence analyses showed differences across the three largest national subgroups. 130 (76.1%) rated ease of use as excellent, and 145 (84.9%) rated clarity of instructions as excellent. 143 participants (83.6%) indicated that they prefer the 5R-STS over a battery of questionnaires.
Conclusions:
In the context of personalized testing, moving toward individually focused precision assessment of patients requires rigorous external validation to ensure the robustness of such applied computational methods. In this external validation, the model demonstrated limited generalization including a systematic underestimation of slower test times and insufficient personalized ULN calibration. These findings indicate that external validity of models derived from single-center data can be limited, underscoring the importance of comprehensive external validation and, potentially, multicenter retraining before clinical implementation.

