Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
Introduction to Language of Pathophysiology l01:25

Introduction to Language of Pathophysiology l

Pathophysiology investigates how biological mechanisms—typically starting at the cellular level—disrupt normal bodily functions. It bridges anatomy and physiology to explain the progression of disease. With this foundation, it is important to understand the following key terms used to describe disease processes: Diagnosis:The process of identifying a disease using clinical evaluation, including signs (objective evidence like rashes), symptoms (subjective experiences like pain), laboratory test...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A gold-standard French-language annotated corpus of oncological entities with ICD-O normalisation.

Scientific data·2026
Same author

Clarifying the relationship between biomedical and health informatics and digital health: expert perspectives.

BMJ health & care informatics·2026
Same author

QT prolongation alerts lead to monitoring but rarely to therapeutic changes: a prospective hospital study.

Frontiers in pharmacology·2026
Same author

Can LLMs Turn French PET/CT Narrative Reports into Structured Knowledge?

Studies in health technology and informatics·2026
Same author

Explainable Framework for Ontology-Based Similarity: A Use Case on SNOMED CT.

Studies in health technology and informatics·2026
Same author

Advancing Knowledge in Evaluating the Clinical Impact of Large Language Models for Clinical Text Summarization: A Narrative Review.

Studies in health technology and informatics·2026

Related Experiment Video

Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Benchmarking Open-Source Large Language Models in Medical French.

Maria Tcherepanova1,2, Amandine Quercia2,3, Nikola Bjelogrlic1

  • 1Division of Medical Information Sciences, Geneva University Hospitals, Geneva, Switzerland.

Studies in Health Technology and Informatics
|May 23, 2026
PubMed
Summary

This study evaluated 15 open-source Large Language Models (LLMs) in medical French, finding performance varied widely. The best models show promise for clinical use, but some inconsistencies remain.

Keywords:
BenchmarkClinical AILLMMedical FrenchNLP

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Related Experiment Videos

Last Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Area of Science:

  • Medical Informatics
  • Natural Language Processing
  • Artificial Intelligence in Healthcare

Background:

  • Large Language Models (LLMs) show potential in healthcare but require evaluation in specific languages and domains.
  • Previous research (MedFrenchmark, 2024) highlighted the need for assessing LLMs in medical French.
  • Limited studies exist on the performance of open-source LLMs in specialized medical contexts.

Purpose of the Study:

  • To evaluate the performance of 15 open-source Large Language Models (LLMs) using a French medical question dataset.
  • To assess factual accuracy, contextual relevance, clarity, and clinical usability of LLM-generated responses.
  • To provide an updated overview of open-source LLM capabilities in medical French.

Main Methods:

  • Utilized a subset of 77 medical questions across various specialties and reasoning types.
  • Manually rated 1,155 generated responses from 15 open-source LLMs on a 0-100 scale.
  • Assessed models based on factual accuracy, contextual relevance, clarity, and clinical usability.

Main Results:

  • Model performance varied significantly, ranging from 36% to 81%.
  • GPT-OSS:120B scored highest, but smaller models like Qwen3:8B and Gemma3:27B showed comparable precision and coherence.
  • Stronger models exhibited improved reasoning and fluency, though terminology and semantic control issues persisted.

Conclusions:

  • Open-source LLMs demonstrate significant progress in medical French applications.
  • Despite advancements, challenges in terminology and semantic control require further attention.
  • The evaluated LLMs show promising potential for future clinical integration in French-speaking healthcare settings.