Related Experiment Video
Updated: Oct 9, 2026

Multi-Modal Signals for Analyzing Pain Responses to Thermal and Electrical Stimuli
Published on: April 5, 2019
Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language
Xiaoyi Zhang1, Christopher R Wilson2, Hannah Eyre2
1Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, United States.
Background:
Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring.
Objective:
The objective of this study is to develop and validate a domain-informed natural language processing framework to derive pain severity and functional interference outcomes from unstructured clinical narratives. We aim to demonstrate that natural language processing-derived outcomes can serve as a reliable, scalable surrogate for resource-intensive patient-reported surveys.
Methods:
This study uses a retrospective cohort of 3725 Veterans with chronic musculoskeletal pain initiating complementary and integrative health therapies across 18 Veterans Health Administration Whole Health Flagship sites (2021-2023). The dataset encompasses longitudinal patient-reported outcome surveys serving as the benchmark, linked with unstructured clinical narratives from the Veterans Health Administration electronic health record. Guided by established psychometric instruments and subject matter expert (SME) input, we developed a seed lexicon and annotation guidelines to identify language distinguishing 3 pain domains: pain severity, interference with enjoyment of life, and interference with general activities. Preliminary large language model (LLM) prompting was used to identify 600 candidate encounters (200 per domain) from 6747 notes across 260 patients for SME annotation, forming a ground truth validation sample. Two candidate LLMs will be evaluated on this sample; the best-performing LLM will generate a large library of span-level annotations to train a scalable, lightweight language model. The study uses a 3-stage validation process: (1) documentation completeness of pain interference in clinical narratives against SME-annotated references; (2) inference accuracy of the LLM-as-annotator and the fine-tuned lightweight model against SME annotations across note-level classification and span-level localization; and (3) concordance between the lightweight model's output and patient-reported pain, enjoyment, and general activity scores across a range of temporal windows.
Results:
As of July 2026, the cohort of 3725 Veterans has been identified and linked to clinical notes. The seed lexicon and annotation guidelines have been developed. Applying a developmental LLM to screen 6747 text notes in 6642 unique encounters over a 7-month period for 260 patients, at least 1 of the 3 pain domains was identified in 75% of notes and 99% of patients. SME validation at the encounter level is in progress. Final results from the subsequent knowledge distillation and validation stages are expected in the first half of 2027.
Conclusions:
This protocol outlines a framework for identifying severe pain intensity and interference from clinical narratives, addressing a critical gap in health care system surveillance. To our knowledge, this is the first study to validate clinical text-based pain outcome extraction against patient-reported outcomes in a nationwide longitudinal cohort. If successful, this approach will enable health care systems to continuously monitor reports of pain-related functional interference and support more holistic, patient-centered pain management at scale.