Related Experiment Video
Updated: Jun 22, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Data Set and Benchmark (MedGPTEval) to Evaluate Responses From Large Language Models in Medicine: Evaluation
Jie Xu1, Lu Lu1, Xinwei Peng1
1Shanghai Artificial Intelligence Laboratory, OpenMedLab, Shanghai, China.
A new evaluation system, MedGPTEval, was developed to assess large language models (LLMs) in medicine. MedGPTEval found that the Dr PJ model demonstrated superior performance in medical dialogues and case reports compared to other LLMs.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in clinical applications but risk generating unreliable responses (hallucinations).
- Hallucinations in medical LLMs pose significant patient safety risks.
- Systematic evaluation of medical LLMs is crucial for identifying and mitigating these risks.
Purpose of the Study:
- To develop and validate a comprehensive evaluation system, MedGPTEval, for assessing large language models in the medical domain.
- To benchmark the performance of leading LLMs, including ChatGPT, ERNIE Bot, and Dr PJ, using the developed system.
Main Methods:
- Designed evaluation criteria through literature review and expert consensus (Delphi method).
- Created Chinese medical datasets, including dialogues and case reports, for LLM interaction.
- Conducted blind evaluations of LLM-generated responses by licensed medical experts using established criteria.
Main Results:
- The developed MedGPTEval system includes 16 indicators across medical, social, contextual, and robustness capabilities.
- Dr PJ demonstrated superior performance over ChatGPT and ERNIE Bot in multi-turn medical dialogues and case report scenarios.
- Dr PJ showed better robustness in terms of semantic consistency and error rates, though ChatGPT slightly outperformed it in medical professional capabilities.
Conclusions:
- MedGPTEval offers a robust framework for evaluating medical LLMs, including open-source datasets and benchmarks.
- Experimental results indicate Dr PJ's superior performance in social and professional medical contexts compared to ChatGPT and ERNIE Bot.
- The MedGPTEval system is readily adoptable by researchers to enhance open-source medical LLM datasets and facilitate further development.
More Related Videos
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018