Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The potential of LLMs in generating questions and answers with EHRs
Yunqi Zhu1,2,3, Wen Tang4, Huayu Yang4
1Guangzhou University, Guangzhou, China.
Background:
This study aimed to generate medical qualification exam questions and their corresponding answers from real-world electronic health records (EHRs) with large language models (LLMs), and to compare their output to that of human medical experts.
Methods:
Utilizing a multicenter bidirectional anonymized database China Elderly Comorbidity Medical Database (CECMed), a total of 8 LLMs: ERNIE 4, ChatGLM 4, Doubao, Hunyuan, Spark 4, Qwen, Llama 3, and Mistral were tasked with generating open-ended questions and answers based on a subset of sampled admission reports. LLMs generated the medical question and answer through few-shot prompting. An independent expert panel scored the AI-generated outputs based on multiple criteria, including coherence, sufficiency of key information, information correctness, factual consistency, evidence of statement, and professionalism, using 5-point Likert scales.
Results:
For question generation, ERNIE 4 achieved the highest cumulative score (16.47). Human experts surpassed LLMs in sufficiency of key information (3.67) but lagged in information correctness (3.63 vs. LLMs' 4.03-4.57). The information correctness of ERNIE was significantly higher than the human's [0.93 (0.62, 1.24), p < 0.01]. For answer generation, humans led overall (14.49), while Doubao outperformed the other LLMs in coherence (3.57), factual consistency (3.60), and professionalism (3.53). The coherence of human's was significantly better than that of 8 LLMs, especially outperformed Llama [0.8 (0.37, 1.23), p < 0.01] and Mistral [0.87 (0.45, 1.28), p < 0.01].
Conclusions:
Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming. This study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians. Although current LLMs performed dissatisfactorily in some aspects, medical students and interns may find LLMs a useful auxiliary tool to support their learning.
Clinical Trial Registration:
https://clinicaltrials.gov/study/NCT06316544, identifier: NCT06316544.
Related Concept Videos
Methods of Documentation VII: EMR
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes:
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Purpose of Health Records II
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include: