Related Experiment Video
Updated: Jun 6, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
504
Analyzing evaluation methods for large language models in the medical field: a scoping review
Junbok Lee1,2, Sungkyung Park3, Jaeyong Shin4,5
1Institute for Innovation in Digital Healthcare, Yonsei University, Seoul, Republic of Korea.
BMC Medical Informatics and Decision Making
|November 29, 2024
Summary
This review of Large Language Models (LLMs) in medicine highlights the need for standardized evaluation frameworks. Current studies show varied methodologies, emphasizing the importance of systematic approaches for future medical LLM research.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Natural Language Processing
Background:
- The rapid rise of Large Language Models (LLMs) necessitates robust evaluation in the medical field.
- Existing performance evaluation studies lack a standardized framework for assessing medical LLM applicability.
Purpose of the Study:
- To systematically review and analyze methodologies used in current LLM evaluations within the medical domain.
- To provide a foundational reference for researchers developing future LLM studies in healthcare.
Main Methods:
- A scoping review of three major databases (PubMed, Embase, MEDLINE) was conducted.
- Articles published between January 1, 2023, and September 30, 2023, were analyzed for evaluation methods, query counts, evaluator types, and prompt engineering use.
Main Results:
- 142 articles met the inclusion criteria, with evaluations primarily using test examinations (37.3%) or medical professional assessments (56.3%).
- Most studies employed fewer than 100 questions, with limited use of repeated measurements (24.2%) and prompt engineering (12.9%).
- Medical assessments predominantly used under 50 queries and involved two evaluators.
Conclusions:
- Further research is crucial for the systematic application of LLMs in healthcare.
- Future studies should prioritize performance improvement and adopt well-structured methodologies.
- Standardized evaluation frameworks are essential for advancing medical LLM research.

