Related Experiment Video
Updated: Jan 24, 2026

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
Published on: April 1, 2016
Comparative evaluation of ChatGPT, Gemini, and Grok in clinical decision-making and general knowledge assessment for
Genta Agani Sabah1, Mehmet Gümüş Kanmaz2
1Department of Orthodontics, Izmir Tinaztepe University, Izmir, Türkiye.
Objective:
This study aimed to compare extraction versus orthodontic eruption decisions for impacted maxillary canines made by three artificial intelligence-based chatbots (ChatGPT, Gemini, and Grok) with those made by orthodontist raters, and to evaluate the overall accuracy of these artificial intelligence-generated recommendations.
Methods:
Thirty-three patients with impacted maxillary canines were selected, and standardized case scenarios incorporating key diagnostic parameters were presented to the three chatbots. Their treatment decisions were recorded and compared with orthodontists' consensus decisions. Additionally, 10 general queries regarding impacted maxillary canines were submitted to the chatbots. The responses were rated by three orthodontists using a modified 5-point Global Quality Score.
Results:
The chatbots and orthodontists showed moderate agreement regarding treatment decisions (κ = 0.411-0.524, P < 0.05). Gemini produced significantly more discordant responses, frequently over-recommending orthodontic eruptions (P = 0.002), whereas Grok and ChatGPT received significantly higher scores than Gemini in the case-based scenarios (P < 0.001). Grok outperformed both ChatGPT and Gemini for general queries (P = 0.006).
Conclusions:
While Gemini showed lower clinical alignment with orthodontists for treatment decisions regarding impacted canines, ChatGPT and Grok demonstrated moderate agreement with orthodontists and produced relatively accurate responses. These findings highlight the potential of chatbots as supportive tools for orthodontic decision-making. However, their use requires careful supervision to avoid the risks associated with inaccurate or misleading recommendations.
Related Concept Videos
Angina III: Clinical Manifestations and Assessment
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Impact of Groups on Groups
Irritable Bowel Syndrome II: Clinical Features and Diagnostic Evaluation
Irritable Bowel Syndrome (IBS) is classified into subtypes based on the predominant bowel habits as determined by the Bristol Stool Form Scale (BSFS). The subtypes are:
Peripheral Arterial Disease II: Clinical Manifestations and Diagnostic Evaluation
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...

