Related Experiment Video
Updated: Feb 2, 2026

Evaluation of Motor Impairment in C. elegans Models of Amyotrophic Lateral Sclerosis
Published on: September 2, 2021
Model-Based and Model-Free Techniques for Amyotrophic Lateral Sclerosis Diagnostic Prediction and Patient Clustering
Ming Tang1,2, Chao Gao1,2, Stephen A Goutman3
1Statistics Online Computational Resource, Department of Health Behavior and Biological Sciences, University of Michigan, Ann Arbor, MI, 48109, USA.
This study used machine learning on a large dataset to stratify patients with Amyotrophic Lateral Sclerosis (ALS) into distinct subgroups, improving understanding of disease progression and aiding clinical decision-making for ALS.
Area of Science:
- Neuroscience
- Data Science
- Biostatistics
Background:
- Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disease affecting approximately 5 in 100,000 individuals in the US.
- Accurate forecasting of ALS progression and patient phenotyping are crucial for clinical decision support.
- Existing data analysis methods have limitations in predicting individual patient trajectories.
Purpose of the Study:
- To develop and compare machine learning models for predicting ALS disease progression, measured by changes in the Amyotrophic Lateral Sclerosis Functional Rating Scale (ALSFRS) score.
- To employ unsupervised learning for reliable and reproducible clustering of ALS patients into distinct computable phenotypes.
- To leverage the PRO-ACT database for robust data-driven insights into ALS patient heterogeneity.
Main Methods:
- Utilized a large ALS dataset (8,000 patients, 3 million records) from the PRO-ACT archive.
- Applied both model-based (linear models) and model-free (random forest, BART) machine learning techniques for outcome prediction.
- Employed unsupervised machine learning for patient cohort clustering and computable phenotype derivation.
Main Results:
- Predicting individual ALSFRS score changes yielded moderate success (correlation coefficients up to 0.545).
- Unsupervised clustering reliably stratified patients into four distinct computable phenotypic subgroups with over 95% consistency.
- Salient clinical features from the PRO-ACT archive were identified to explicate the derived patient subcohorts.
Conclusions:
- While direct prediction of univariate clinical outcomes in ALS remains challenging, data science strategies effectively cluster patients.
- Unsupervised clustering generates stable and reliable computable phenotypes, offering valuable insights into ALS patient heterogeneity.
- These findings support the generation of evidence-based hypotheses regarding complex, multivariate factors influencing ALS progression.
Related Concept Videos
Stereotype Content Model
Predicting Molecular Geometry
Molecular Models
Lateralization
The Bohr Model
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...

