Related Experiment Video
Updated: Aug 5, 2026

Multi-Modal Signals for Analyzing Pain Responses to Thermal and Electrical Stimuli
Published on: April 5, 2019
Assessing Pain Catastrophizing Through Free-Text Responses: A Validation of Large Language Models
Angela Lee1, Dokyoung Sophia You2, Troy C Dildine1
1Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University School of Medicine, Stanford University, 1070 Arastradero Road, Suite 200, MC5596, Palo Alto, CA, 94304, USA.
None:
Validated measures of pain catastrophizing primarily assess catastrophizing as a stable trait. However, emerging evidence suggests catastrophizing fluctuates with context, highlighting a need for ecologically valid methods to capture it. This study evaluated large language models (LLMs) as implicit markers of catastrophizing from free-text responses from ninety-one adults with chronic pain receiving long-term opioid therapy (57.3% Female; mean age = 60.5 years). Patients completed baseline measures, including the trait pain catastrophizing scale (PCS), followed by a 10-minute writing task after random assignment to a negative, positive, or neutral pain-coping condition. State affect and pain were assessed before and after writing tasks and again after a cold pressor task (4°C; ≤ 2 minutes). A state PCS followed the cold pressor task. Free-text responses were analyzed using four LLMs (Claude Opus 4; GPT Mini 4o; Llama 4 Maverick; and Gemini 2.5 Pro). ANOVA-based results supported discriminant validity, as all four LLM-derived pain catastrophizing scores differentiated negative from positive and neutral pain-coping conditions. Convergent validity was model-dependent; only Gemini-derived scores correlated with state catastrophizing (r = .22) and pain unpleasantness (r = .23). Divergent validity was mixed. LLM-derived scores were unrelated to pain intensity, but Gemini and Claude-derived scores showed small correlations with trait PCS (r's = .21; 28, respectively). All LLM-derived scores also correlated with negative affect (r's = .29-.41), comparable in magnitude to state PCS, suggesting limited specificity. These findings provide preliminary evidence that certain LLMs may serve as implicit markers of state pain catastrophizing, but further study is needed.

