Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Arabian Nights or English Days? Accuracy of Large Language Models in Answering Bilingual Dental Multiple-Choice
Hala Alanazi1, Lujain Altalhi1, Nadeen Alanazi1
1King Abdullah International Medical Research Center, College of Dentistry, King Saud Bin Abdulaziz University for Health Sciences, Ministry of National Guard Health Affairs, Riyadh, Kingdom of Saudi Arabia.
Background:
While large language models (LLMs) perform well in medical education, their ability to accurately interpret and answer English and Arabic dental multiple-choice questions (MCQs) remains underexplored.
Aims:
This study evaluated the performance of advanced LLMs in answering dental MCQs in both languages, identifying language-specific challenges and assessing their applicability in multilingual dental education.
Materials And Methods:
A total of 300 MCQs from ten dental specialties were sourced from question banks. The MCQs were translated into Arabic and reviewed for linguistic and technical accuracy. Four LLMs, ChatGPT-4o, ChatGPT-4, Gemini, and Claude, were tested separately on Arabic and English datasets. Accuracy was the primary metric, along with specialty-specific performance, question type differentiation, and cross-language consistency.
Results:
Claude achieved the highest accuracy in English (89%), while Gemini performed best in Arabic (80%). Most models showed better performance in English, with notable translation inconsistencies, particularly for ChatGPT models. Specialty-wise, Claude and Gemini excelled in endodontics and operative dentistry. No significant differences were observed between knowledge-based and clinical questions, but Arabic interpretation posed challenges. Statistical analysis confirmed significant differences between models and across languages.
Discussion:
Gemini demonstrated robust performance in Arabic, while Claude excelled in English. ChatGPT models exhibited limitations, particularly in Arabic datasets. Performance varied across specialties, highlighting the need for improved multilingual adaptability and specialty-specific training.
Conclusion:
Expanding specialised and culturally relevant datasets is essential for optimising LLMs' educational utility. This study provides key insights into LLM performance in bilingual dental education, supporting future advancements in AI-driven learning tools.

