Related Experiment Video
Updated: May 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation and Bias Analysis of Large Language Models in Generating Synthetic Electronic Health Records: Comparative
Ruochen Huang1, Honghan Wu2, Yuhan Yuan1
1School of Biomedical Engineering and Informatics, Nanjing Medical University, Nanjing, China.
Large language models (LLMs) generate synthetic electronic health records (EHRs) with improved completeness but amplified gender and racial biases. Addressing these performance-bias trade-offs is crucial for equitable healthcare AI.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Health Equity
Background:
- Synthetic electronic health records (EHRs) generated by large language models (LLMs) offer privacy-preserving solutions for clinical education and model training.
- However, underexplored performance variations and demographic biases in LLM-generated EHRs pose risks to equitable healthcare.
Purpose of the Study:
- To systematically assess the performance of various LLMs in generating synthetic EHRs.
- To critically evaluate gender and racial biases in LLM-generated EHRs across 20 diseases with varying demographic prevalence.
Main Methods:
- Developed a framework to generate 140,000 synthetic EHRs using 7 LLMs and 10 prompts.
- Introduced the electronic health record performance score (EPS) for completeness and statistical parity difference (SPD) for demographic bias assessment.
- Utilized chi-square tests to evaluate bias across demographic groups.
Main Results:
- Larger LLMs demonstrated superior EHR generation performance (higher EPS) but exhibited increased gender and racial biases.
- Observed sex polarization, with amplified representation for female-dominated diseases and skewed male representation in others.
- Identified racial biases, including overestimation of White/Black populations and underestimation of Hispanic/Asian groups in synthetic EHRs.
Conclusions:
- A performance-bias trade-off exists: larger LLMs generate more comprehensive EHRs but with heightened demographic biases.
- Biases were present across all tested LLMs, not exclusively larger models.
- Findings underscore the urgent need for bias mitigation strategies and fairness benchmarks in healthcare AI.
More Related Videos
Related Concept Videos
Methods of Documentation VII: EMR
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes:
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Bias in Epidemiological Studies
Mechanistic Models: Compartment Models in Individual and Population Analysis

