Related Experiment Video
Updated: Aug 11, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Comparative Assessment of Performance of Domain-Specific and Primed Versus Non-Primed Artificial Intelligence
Shajaratul Yakeen Nabi1, Aaqib Shah2, Hasibe Elif Kuru3
1Department of Pediatric and Preventive Dentistry, Government Dental College and Hospital, Srinagar, Jammu & Kashmir, India.
Background/Aims:
To compare the performance of a domain-specific dental trauma chatbot with general-purpose artificial intelligence chatbots under primed and non-primed conditions for clinical decision-making in injuries to the primary dentition.
Methods:
Fifteen standardized clinical case scenarios were developed for evaluating three chatbot systems: a domain-specific model (Dental Trauma Evo) and two general-purpose models (ChatGPT 5.2 and Perplexity Pro). General-purpose chatbots were evaluated under primed and non-primed conditions, where priming involved providing a summarized guideline document prior to scenario input. Chatbots were required to generate responses addressing diagnosis, immediate management, follow-up intervals, and radiographic recommendations. Two Pediatric dentists independently evaluated responses using a binary scoring system based on International Association of Dental Traumatology (IADT) guidelines. Statistical comparisons were performed using McNemar's test and agreement analysis with kappa statistics.
Results:
All chatbots demonstrated complete diagnostic accuracy across scenarios. The domain-specific chatbot achieved 100% accuracy in immediate management decisions, while minor inaccuracies were observed among general-purpose models. The most substantial differences were observed in follow-up interval and radiographic recommendations. Priming significantly improved the performance of general-purpose chatbots, with ChatGPT and Perplexity demonstrating marked gains in follow-up interval accuracy and radiographic recommendations. In these domains, primed general-purpose models demonstrated greater agreement with IADT guideline recommendations than the domain-specific chatbot.
Conclusions:
Guideline-based priming was associated with improved performance of general-purpose AI chatbots in clinical decision-making for injuries to the primary dentition, particularly for follow-up scheduling and radiographic recommendations. While domain-specific models provide reliable management guidance, contextual exposure to clinical guidelines enables general large language models to demonstrate a high level of agreement with established recommendations in certain aspects of trauma care in primary dentition. These findings highlight the potential role of prompt-guided AI systems as supportive tools in dental traumatology decision-making.