Related Experiment Video
Updated: Aug 12, 2026

Mechanical Ventilation Boot Camp Curriculum
Published on: March 12, 2018
Comparative benchmarking of three artificial intelligence chatbots (ChatGPT-5.1, Qwen- 3 Max, and Perplexity AI) on
Khouloud Kchaou1, Salma Mokaddem1, Soumaya Khaldi1
1Laboratory of Physiology, Faculty of Medicine of Tunis, University of Tunis El Manar, Tunis, Tunisia.
Background:
Artificial intelligence (AI) chatbots are increasingly used by medical students for learning and examination preparation. However, their reliability in mechanistically demanding disciplines such as respiratory physiology remains insufficiently evaluated using authentic examination material. This study aimed to assess and compare the performance of three AI chatbots on validated undergraduate respiratory physiology examination questions.
Methods:
We conducted a cross-sectional analytical study including all respiratory physiology multiple-choice questions (MCQs), each comprising five propositions, used in official undergraduate examinations at the Faculty of Medicine of Tunis during the academic years 2023-2024 and 2024-2025 (101 questions). Three AI chatbots (ChatGPT-5.1, Qwen-3 Max, and Perplexity AI) were evaluated using standardized French-language prompts. Exact question-level concordance with the faculty-validated answer key, established by the teaching staff responsible for respiratory physiology examinations, was compared using Cochran's Q test and pairwise McNemar tests. At the proposition level, responses were analyzed as binary outcomes (true/false) to compute accuracy, sensitivity, and specificity.
Results:
Question-level exact concordance differed significantly across models (Cochran's Q = 10.18, p = 0.006). ChatGPT showed the highest concordance rate (79.2%), followed by Qwen (74.3%) and Perplexity (62.4%). At the proposition level (505 propositions), ChatGPT achieved the highest overall accuracy (94.1%) and sensitivity for true propositions (96.5%), whereas Qwen demonstrated the highest specificity for false propositions (94.2%).
Conclusion:
Although all evaluated AI chatbots performed well on respiratory physiology MCQs, substantial variability was observed across models. These findings indicate that AI chatbots are not interchangeable and underscore the need for institution-level evaluation using local examination material before their integration into medical education.
