Related Experiment Video
Updated: Jun 3, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
Assessment of LLMs for Generating Synthetic Health-Related Quality-of-Life Survey Responses Among U.S. Adults
Seungkook Baek1, Sunghyun Lee1, Hyungjin Kim1
1intellicia, Seoripul-gil 29-7, Seocho-gu, Seoul, Republic of Korea.
Introduction:
Large Language Models (LLMs) present a potential alternative to traditional surveys by generating synthetic health-related quality-of-life (HRQoL) survey responses, offering a faster and more cost-effective approach to population health monitoring. This study evaluated the accuracy of LLM-generated predictions of population-level HRQoL survey response.
Methods:
The primary outcomes were the Physical Component Summary (PCS) and Mental Component Score (MCS) scores, calculated from responses to the 12-item Short Form Survey (SF-12). LLaMA 4 was used to generate synthetic item-level SF-12 responses; PCS and MCS scores were then derived using the standard SF-12 scoring algorithm. Prompts were constructed in four scenarios: no individual-level characteristics, demographic characteristics, demographic and socioeconomic characteristics, and demographic, socioeconomic, and health-related characteristics. Data from the 2022 Medical Expenditure Panel Survey were analyzed in 2025.
Results:
LLM performance in estimating PCS was poor when no individual-level data were provided (mean absolute percentage error [MAPE]: 11.10; root mean square error [RMSE]: 156.53; R²: 0.00; prediction-to-observation ratio: 0.97; Pearson correlation: 0.02). Accuracy improved substantially with the addition of individual-level information, achieving the best performance when demographic, socioeconomic, and health-related characteristics were all incorporated (MAPE: 7.91; RMSE: 109.17; R²: 0.26; prediction-to-observation ratio: 1.00; Pearson correlation: 0.51). However, limitations persisted, particularly in accurately capturing values at the distributional extremes. For MCS, inclusion of individual characteristics led to only modest improvements.
Conclusions:
LLMs showed potential to generate synthetic SF-12 response patterns, particularly for physical HRQoL; however, limitations in individual-level accuracy and performance at the distributional extremes underscore the need for further methodological refinement.
Related Concept Videos
Data Collection by Survey
Surveys
Longitudinal Research
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
Dimensions of Health and Illness
Health Literacy
