Data Set and Benchmark (MedGPTEval) to Evaluate Responses From Large Language Models in Medicine: Evaluation

Jie Xu1, Lu Lu1, Xinwei Peng1

  • 1Shanghai Artificial Intelligence Laboratory, OpenMedLab, Shanghai, China.

PubMed
Summary

A new evaluation system, MedGPTEval, was developed to assess large language models (LLMs) in medicine. MedGPTEval found that the Dr PJ model demonstrated superior performance in medical dialogues and case reports compared to other LLMs.

Related Concept Videos