Related Experiment Video
Updated: Mar 1, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Do Large Language Models Perform Equally Across Languages? A Comparison of Responses to Frequently Asked Questions in
Hadi Ufuk Yörükoğlu1, Can Aksu2, Pervez Sultan3
1Department of Anesthesiology and Reanimation, Kocaeli University School of Medicine, Kocaeli, Turkey.
ChatGPT 4.0 and DeepSeek V3 large language model (LLM) chatbots provide superior English responses compared to Turkish ones in anesthesia patient education. ChatGPT 4.0 English responses were better than DeepSeek V3.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Anesthesiology
Background:
- Large language model (LLM) chatbots are increasingly used in healthcare for patient education.
- Evaluating the reliability and understandability of LLM responses in multiple languages is crucial, especially in anesthesia.
- Patient education in anesthesia requires accurate and accessible information.
Purpose of the Study:
- To compare the quality of English responses from ChatGPT 4.0 and DeepSeek V3.
- To evaluate content and communication differences between English and Turkish responses from these LLMs.
- To assess the suitability of LLMs for anesthesia patient education.
Main Methods:
- Anesthesiologists proficient in English and Turkish served as expert evaluators.
- Ten frequently asked anesthesia questions were selected and translated.
- Responses from ChatGPT 4.0 and DeepSeek V3 in both languages were assessed for overall, content, and communication quality.
Main Results:
- ChatGPT 4.0's English responses were significantly superior to DeepSeek V3's English responses (P<0.001).
- Both ChatGPT 4.0 and DeepSeek V3 demonstrated significantly better overall, content, and communication quality in English compared to their Turkish responses (P<0.001).
- English responses from both LLMs outperformed their Turkish counterparts across all evaluated metrics.
Conclusions:
- ChatGPT 4.0 exhibits higher overall response quality than DeepSeek V3 in English for anesthesia-related queries.
- English language responses from both evaluated LLMs are superior to their Turkish counterparts for anesthesia patient education.
- Further research is needed to optimize LLM performance in non-English languages for healthcare applications.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
13:12Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...