Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Utilizing Large Language Models to Enhance Patient-Reported Outcome Measures: Application to the EQ-5D-5L and
Jan M Heijdra Suasnabar1, Marieke van Buchem2, Mathieu F Jansen3
1Department of Biomedical Data Science, Leiden University Medical Center, the Netherlands.
Summary
Large language models (LLMs) show promise in developing patient-reported outcome measures (PROMs). This study used LLMs to identify potential EQ-5D-5L dimensions from patient text data, demonstrating their utility in PROM development.
Area of Science:
- Health Services Research
- Medical Informatics
- Psychometrics
Background:
- Patient-reported outcome measures (PROMs) are crucial for evaluating health interventions.
- Developing and adapting PROMs, such as the EQ-5D-5L, requires robust methods to capture patient experiences.
- Large language models (LLMs) offer novel computational approaches for analyzing qualitative patient data.
Purpose of the Study:
- To evaluate the feasibility of using LLMs to identify potential bolt-on dimensions for the EQ-5D-5L.
- To assess the performance of LLMs in analyzing patient-reported free-text data for PROM development.
- To compare LLM-derived dimensions with those identified through traditional qualitative and topic modeling approaches.
Main Methods:
- GPT-4o was employed to analyze free-text narratives from 1,977 celiac disease patients.
- Prompts were designed to elicit potential EQ-5D-5L bolt-on dimensions and draft item wordings.
- LLM-identified dimensions were compared against qualitative analysis and topic modeling results using Kappa statistics.
- Suitability of LLM-generated item wordings was assessed against established criteria.
Main Results:
- The LLM identified 12 potential bolt-on dimensions, with 9 overlapping with qualitative analysis and 5 with topic modeling.
- Text-entry level agreement between LLM and qualitative methods was generally moderate to almost perfect (median Kappa=0.68).
- LLM-generated item wordings for the top 4 dimensions scored highly (4.0-4.4/5) on suitability assessments.
Conclusions:
- LLMs demonstrate significant potential as tools to aid in the development and modification of PROMs using patient-generated text.
- The findings support the use of LLMs for identifying novel dimensions and generating preliminary item content for PROMs.
- Further research is recommended to explore the transferability of this LLM-based approach across different disease areas and data sources.