Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

MedNLP-Hub: A knowledgebase platform for biomedical NLP tool discovery in clinical informatics.

International journal of medical informatics·2026
Same author

Patient Selection Metrics and Efficacy of Neoadjuvant Chemotherapy in Advanced Epithelial Ovarian Cancer: A Retrospective Analysis.

International journal of women's health·2026
Same author

Augmenting large language models with clinical knowledge graph for personalized perioperative fluid therapy question answering.

PLOS digital health·2026
Same author

PEPRKD-depression: A knowledge database supporting evidence-based personalized exercise prescription recommendations in depression.

Digital health·2026
Same author

Microenvironment T-Type calcium channels regulate neuronal and glial processes in tumor cells to promote glioblastoma growth.

Neuro-oncology·2026
Same author

Benchmarking Large Language Models and Prompt Engineering Strategies in Microsatellite Instability Cancers: Evaluation Study.

Journal of medical Internet research·2026

Related Experiment Video

Updated: Apr 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

A Dataset for Evaluating Large Language Models on Chinese National Medical Licensing Examinations.

Hui Zong1, Jiaxue Cha2, Jiao Wang1

  • 1Joint Laboratory of Artificial Intelligence for Critical Care Medicine, Department of Critical Care Medicine and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital, Sichuan University, Chengdu, 610041, China.

Scientific Data
|April 17, 2026
PubMed
Summary

A new benchmark dataset, CNMLEQA, evaluates large language models (LLMs) on the Chinese National Medical Licensing Examination. This dataset helps advance AI in Chinese medical education and clinical reasoning.

More Related Videos

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

3.4K
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
08:32

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks

Published on: September 5, 2019

6.0K

Related Experiment Videos

Last Updated: Apr 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

3.4K
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
08:32

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks

Published on: September 5, 2019

6.0K

Area of Science:

  • Artificial Intelligence
  • Medical Informatics
  • Natural Language Processing

Background:

  • Large language models (LLMs) show promise in medical education and clinical reasoning.
  • A lack of standardized, non-English datasets hinders LLM evaluation in diverse medical contexts.

Purpose of the Study:

  • Introduce CNMLEQA, a novel benchmark dataset for evaluating LLMs on the Chinese National Medical Licensing Examination.
  • Address the gap in non-English medical datasets for AI model assessment.

Main Methods:

  • CNMLEQA dataset creation by integrating question-answer pairs from PubMed, GitHub, and MedExamLLM.
  • Dataset comprises two subsets (CNMLEQA-10k and CNMLEQA-3k) with multiple-choice questions.
  • Questions annotated by clinical experts across dimensions: type, year, and clinical scenarios (disease, surgery, medication, lab, symptom).

Main Results:

  • Evaluated state-of-the-art LLMs including Gemini, DeepSeek, GPT, Qwen, and LLaMA.
  • Qwen2.5-32B achieved 90.88% accuracy on CNMLEQA-10k; DeepSeek-R1 achieved 91.59% on CNMLEQA-3k.
  • Fine-tuning experiments demonstrated significant performance improvements for Qwen models.

Conclusions:

  • CNMLEQA offers a multidimensional, clinically grounded benchmark for LLM evaluation in Chinese medical applications.
  • The dataset facilitates the development and assessment of AI tools for Chinese medical education and practice.