Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Medical Text Summaries Using Automatic Evaluation Metrics and LLM-as-a-Judge Approach: A Pilot Study
Yuriy Vasilev1, Irina Raznitsyna1, Anastasia Pamova1,2
1Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies of the Moscow Health Care Department, 127051 Moscow, Russia.
Large Language Models (LLMs) show promise for summarizing electronic health records (EHRs). However, automated quality control methods, including LLM-as-a-judge, struggle to detect factual errors, necessitating expert review.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
Background:
- Electronic health records (EHRs) contain vital clinical data but are challenging to process.
- Large Language Models (LLMs) offer a promising solution for summarizing EHR data to aid physicians.
- Automated quality control is essential for integrating LLM summarization tools into clinical practice.
Purpose of the Study:
- To assess the feasibility and limitations of automated quality control for LLM-generated medical summaries.
- To evaluate automatic metrics and LLM-as-a-judge approaches without expert involvement.
Main Methods:
- Six open-source LLMs generated summaries from 30 EHR text samples.
- Summaries were evaluated using standard metrics (BLEU, ROUGE, METEOR, BERTScore) and LLM-as-a-judge.
- Criteria included relevance, completeness, redundancy, coherence, grammar, terminology, and hallucination detection.
- Expert evaluation was performed using the same criteria for comparison.
Main Results:
- LLMs demonstrate significant potential for summarizing medical data.
- Neither automatic metrics nor LLM judges reliably detect factual errors or semantic distortions (hallucinations).
- A Pearson correlation of 0.688 was observed between LLM summary quality scores and expert opinions regarding relevance.
Conclusions:
- Fully automating the quality evaluation of medical summaries remains a significant challenge.
- Future research should prioritize hallucination detection methods and explore larger, specialized LLMs for medical text.
- Integrating retrieval-augmented generation (RAG) into LLM-as-a-judge architectures warrants further investigation.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019