Related Experiment Video
Updated: Jun 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review
Suhana Bedi1, Yutong Liu2, Lucy Orr-Ewing2
1Department of Biomedical Data Science, Stanford School of Medicine, Stanford, California.
Current evaluations of large language models (LLMs) in healthcare primarily assess accuracy on exams, neglecting real patient data and critical aspects like fairness. Future research needs broader scope and clinical relevance.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing (NLP) Applications
- Medical Informatics
Background:
- Large language models (LLMs) show potential for healthcare applications.
- Existing evaluation methods may not fully capture the utility of LLMs in clinical settings.
- A comprehensive understanding of current LLM evaluations is needed to guide future research.
Purpose of the Study:
- To systematically review and categorize existing evaluations of LLMs in healthcare.
- To analyze evaluations based on data type, healthcare task, NLP/NLU tasks, evaluation dimensions, and medical specialty.
- To identify gaps and limitations in current LLM evaluation strategies within the medical domain.
Main Methods:
- Systematic literature search of PubMed and Web of Science (January 2022 - February 2024).
- Inclusion of studies evaluating one or more LLMs in healthcare.
- Categorization of 519 studies by independent reviewers across five key components.
Main Results:
- Only 5% of studies utilized real patient care data; most focused on medical knowledge assessment (e.g., licensing exams) and diagnosis.
- Question answering was the predominant NLP/NLU task (84.2%), with limited focus on summarization or dialogue.
- Accuracy was the primary evaluation metric (95.4%), while fairness, bias, toxicity, and deployment considerations were infrequently assessed.
Conclusions:
- Current LLM evaluations in healthcare are heavily skewed towards accuracy on standardized tests, lacking real-world clinical data and diverse task considerations.
- Critical aspects like fairness, bias, and deployment readiness are under-evaluated.
- Future research should prioritize standardized metrics, clinical data utilization, and a broader range of healthcare tasks and specialties for robust LLM assessment.
More Related Videos
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Statistical Software for Data Analysis and Clinical Trials
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes:

