Related Experiment Video
Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Open-Source Large Language Models in Medical French.
Maria Tcherepanova1,2, Amandine Quercia2,3, Nikola Bjelogrlic1
1Division of Medical Information Sciences, Geneva University Hospitals, Geneva, Switzerland.
This study evaluated 15 open-source Large Language Models (LLMs) in medical French, finding performance varied widely. The best models show promise for clinical use, but some inconsistencies remain.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Artificial Intelligence in Healthcare
Background:
- Large Language Models (LLMs) show potential in healthcare but require evaluation in specific languages and domains.
- Previous research (MedFrenchmark, 2024) highlighted the need for assessing LLMs in medical French.
- Limited studies exist on the performance of open-source LLMs in specialized medical contexts.
Purpose of the Study:
- To evaluate the performance of 15 open-source Large Language Models (LLMs) using a French medical question dataset.
- To assess factual accuracy, contextual relevance, clarity, and clinical usability of LLM-generated responses.
- To provide an updated overview of open-source LLM capabilities in medical French.
Main Methods:
- Utilized a subset of 77 medical questions across various specialties and reasoning types.
- Manually rated 1,155 generated responses from 15 open-source LLMs on a 0-100 scale.
- Assessed models based on factual accuracy, contextual relevance, clarity, and clinical usability.
Main Results:
- Model performance varied significantly, ranging from 36% to 81%.
- GPT-OSS:120B scored highest, but smaller models like Qwen3:8B and Gemma3:27B showed comparable precision and coherence.
- Stronger models exhibited improved reasoning and fluency, though terminology and semantic control issues persisted.
Conclusions:
- Open-source LLMs demonstrate significant progress in medical French applications.
- Despite advancements, challenges in terminology and semantic control require further attention.
- The evaluated LLMs show promising potential for future clinical integration in French-speaking healthcare settings.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Statistical Software for Data Analysis and Clinical Trials
Introduction to Language of Pathophysiology l
