Related Experiment Video
Updated: Jun 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging large language models in patient-reported outcome measure development: practical opportunities, cautions,
Constance Mara1,2, Tiffany Rybak3,4
1Behavioral Medicine & Clinical Psychology, Cincinnati Children's Hospital Medical Center, Cincinnati, USA. Constance.Mara@cchmc.org.
Background:
Patient-reported outcome (PRO) measure development involves multiple language-intensive stages, including conceptual domain definition, candidate item generation, qualitative refinement, and cognitive interviewing prior to psychometric evaluation. Large language models (LLMs) may offer new opportunities to support these qualitative development activities, although their role within established PRO development frameworks remains incompletely defined.
Aims:
This commentary proposes a practical human-in-the-loop roadmap for integrating LLMs into the qualitative phases of PRO development while preserving established standards for content validity and psychometric rigor.
Approach:
Drawing on examples from development of the Eating Behavior Measurement (EBM) project and emerging literature from PRO science, survey methodology, and psychological measurement, we outline several bounded use cases for LLMs in PRO measure qualitative development workflows. These include accelerating synthesis of legacy item pools ("domain cartography"), generating candidate items within human-defined constructs, supporting iterative item revision in response to cognitive interview findings, conducting semantic coherence checks for construct alignment, and assisting with developmental or contextual adaptation of candidate items. Across these applications, LLMs function as structured drafting and analytic tools rather than arbiters of validity. We additionally discuss practical risks involving confirmation bias, semantic circularity, transparency, reproducibility, and construct drift, along with strategies for mitigation through human oversight and model triangulation.
Conclusions:
LLMs do not replace qualitative inquiry, expert judgment, or empirical psychometric validation. Rather, they may help support more systematic and scalable qualitative development workflows when used within bounded, human-centered measurement frameworks. The central challenge for PRO science is not whether to adopt these tools, but how to integrate them responsibly without compromising the evidentiary standards on which the field depends.
Related Concept Videos
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care evaluation by...
Introduction to Language of Pathophysiology ll
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...
Improving Translational Accuracy
Improving Translational Accuracy
