Related Experiment Video
Updated: Aug 5, 2026

Treating Low Back Pain in Failed Back Surgery Patients with Multicolumn-lead Spinal Cord Stimulation
Published on: June 26, 2018
Proof of Concept of Natural Language Processing in Predicting Patient-Reported Outcomes in Spinal Cord Stimulation
Mahir Kabir1, Theresa Medina2, Sohail Rajesh Daulat2
1Department of Neuroscience, University of Arizona College of Science, Tucson, AZ, US.
Objectives:
We aim to analyze the use of natural language processing (NLP) models to predict the improvement of patient reported outcome measures (PROMs).
Materials And Methods:
We examined PROMS and clinic notes from 20 patients with spinal cord stimulation obtained preoperatively and postoperatively at one year. PROMS included numeric rating scale (NRS), patient global impression of change (PGIC), Oswestry Disability Index (ODI), Beck Depression Inventory (BDI), McGill Pain Questionnaire (MPQ), and Pain Catastrophizing Scale (PCS). Three text embedding models were used: Term-frequency-inverse document frequency (TF-IDF), BERT-base, and Bio-ClinicalBERT. Positive response was considered ≥50% improvement in NRS and achieving validated minimal clinically important difference (MCID) for other PROMs. Spearman rank correlations (ρ), Area Under Curve, accuracy, sensitivity, and specificity were used to evaluate the NLPs' predictions and classifications of improvement.
Results:
Our cohort consisted of 12 women and eight men with a mean age of 56 years. Of the three models, TF-IDF showed robust rank-order prediction of MPQ (ρ = -0.97, p < 0.001), PCS (ρ = -0.96, p < 0.001), NRS (ρ = -0.76, p < 0.001), ODI (ρ = -0.58, p = 0.007), and BDI (ρ = -0.71, p < 0.001). TF-IDF showed the strongest performance in classifying binary responders on the basis of MCID classifications for PGIC, MPQ, and PCS (AUC 1.0). BERT-base and Bio-ClinicalBERT exhibited greater performance variability, with higher accuracies for all PROMs but PGIC and ODI.
Conclusions:
The order of improvement can be approximated in this specific data set, but identifying individual responders remains elusive. We anticipate with the advancement of machine learning algorithms in clinical settings, the value of such models will continue to increase.
