Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A triaxial benchmark for assessing responses from large language models in traditional Chinese medicine.

Communications medicine·2026
Same author

Machine Learning for Preoperative Assessment and Postoperative Prediction in Cervical Cancer: Multicenter Retrospective Model Integrating MRI and Clinicopathological Data.

JMIR cancer·2025
Same author

Methodological conduct and risk of bias in studies on prenatal birthweight prediction models using machine learning techniques: a systematic review.

BMC pregnancy and childbirth·2025
Same author

Fetal Birth Weight Prediction in the Third Trimester: Retrospective Cohort Study and Development of an Ensemble Model.

JMIR pediatrics and parenting·2025
Same author

Evolution of the "Internet Plus Health Care" Mode Enabled by Artificial Intelligence: Development and Application of an Outpatient Triage System.

Journal of medical Internet research·2024
Same author

Data Set and Benchmark (MedGPTEval) to Evaluate Responses From Large Language Models in Medicine: Evaluation Development and Validation.

JMIR medical informatics·2024

Related Experiment Video

Updated: Jul 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Performance Gaps and Optimization Strategies in Chinese Medical Large Language Models Based on MedBench: Evaluation

Luyi Jiang1, Jiayuan Chen2, Lu Lu2

  • 1Shanghai Health Development Research Center (Shanghai Medical Information Center), Shanghai Institute of Infectious Disease and Biosecurity, Fudan University, Shanghai, China.

JMIR AI
|July 15, 2026
PubMed
Summary

A new error taxonomy reveals significant weaknesses in medical large language models (LLMs), particularly in critical reasoning tasks. This study provides a roadmap to improve LLM accuracy, safety, and trustworthiness in healthcare.

Keywords:
MedBenchartificial intelligencebenchmarkingincorrect responsesmedical LLMsmedical large language modelsperformance optimization

Related Experiment Videos

Last Updated: Jul 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Area of Science:

  • Artificial Intelligence in Medicine
  • Natural Language Processing
  • Medical Informatics

Background:

  • Medical large language models (LLMs) require robust evaluation frameworks for safe and accurate clinical deployment.
  • Current evaluation methods are insufficient for identifying domain-specific errors and cross-modal challenges in medical LLMs.

Purpose of the Study:

  • To develop and apply a granular error taxonomy for systematically identifying weaknesses in leading medical LLMs.
  • To create an actionable roadmap for enhancing the clinical robustness, safety, and trustworthiness of medical AI.

Main Methods:

  • A granular error taxonomy was developed by analyzing 42,766 responses from 10 top medical LLMs on the MedBench benchmark.
  • Incorrect responses were categorized into 8 types: omissions, hallucination, format mismatch, causal reasoning deficiency, contextual inconsistency, unanswered, output error, and deficiency in medical language generation.
  • A 4-level tiered optimization strategy was proposed, including prompt engineering, knowledge-augmented retrieval, hybrid neuro-symbolic architectures, and causal reasoning frameworks.

Main Results:

  • Analysis of 10 leading medical LLMs revealed significant vulnerabilities despite a 0.86 accuracy in medical knowledge recall.
  • Omissions were the most prevalent error type, with a 96.3% omission rate in critical reasoning tasks.
  • Safety and ethics evaluations showed low robustness (0.79) under option-shuffled conditions, indicating systemic weaknesses in knowledge boundary enforcement and multistep reasoning.

Conclusions:

  • The developed granular error taxonomy provides critical insights into medical LLM vulnerabilities.
  • This work establishes a roadmap for enhancing the clinical robustness and trustworthiness of AI in healthcare.
  • The findings redefine evaluation paradigms, promoting the responsible application of AI in high-stakes medical environments.