Related Experiment Video
Updated: Jul 7, 2026

Induction of Periodontitis via a Combination of Ligature and Lipopolysaccharide Injection in a Rat Model
Published on: February 17, 2023
Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A
Yaniv Mayer1,2, Bertha Demetriou2, Giulio Rasperini3
1The Ruth and Bruce Rappaport Faculty of Medicine, Technion, Israel Institute of Technology, Haifa, Israel.
Aim:
To compare retrieval-augmented systems with general-purpose large language models (LLMs) on standardised periodontal clinical vignettes.
Materials And Methods:
Eleven AI systems were evaluated: nine general-purpose LLMs, one general-purpose retrieval-augmented platform (Perplexity) and one medical-domain retrieval-augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5-point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover-Iman tests with Holm correction were applied; mixed-effects and ordinal models served as sensitivity analyses.
Results:
At least one parameter scored dangerous (≤ 2) in 3.3%-46.7% of responses across platforms, despite mean composite scores (3.28-4.86) exceeding the rubric midpoint of 3.0. Between-model differences were significant (p < 0.001), with a small-to-medium overall effect (Kendall's W = 0.17) and large within-category effects (W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier.
Conclusion:
Retrieval-augmented systems rated highest, but this advantage was confounded with response length. The dangerous-response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.
