Related Experiment Video
Updated: Aug 5, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Curated retrieval-augmented generation for dental traumatology
José Espona1, Elena Roig2, Marc Garcia3
1Department of Restorative Dentistry, Universitat Internacional de Catalunya, Barcelona, Spain.
Objectives:
To validate DT-RAG, a curated retrieval-augmented generation system for dental traumatology decision support, against eight commercial large language models.
Methods:
A knowledge base of 250 curated text units ("chunks") from five authoritative sources (IADT 2020, ESE 2021, Krastl 2021, AAE 2013, Cochrane) was built and deployed with Gemini 2.5 Flash as base model. In Study 1, 99 binary clinical questions were submitted in three runs to DT-RAG and eight LLMs; modal accuracy was compared by McNemar exact tests with Holm correction. In Study 2, seven blinded specialists scored DT-RAG against three frontier LLMs on ten clinical scenarios using a 92-point rubric; differences were estimated by linear mixed-effects regression.
Results:
DT-RAG achieved 96.0% modal accuracy (95% CI 90.1-98.4), significantly exceeding every commercial LLM (best comparator GPT-5.5 87.9%; paired difference +8.1 pp, 95% CI +2.9 to +14.9; p_Holm = 0.013). The curated knowledge base elevated the base model from 49.5% to 98.0% valid rationale rate, eliminating confabulations in this evaluation (0 vs 21). In Study 2, DT-RAG achieved the highest mean score (82.1/92; 89.3%), significantly exceeding Claude Opus 4.5 (72.6; paired difference +9.6 points, exact Wilcoxon p = 0.016), Gemini 2.5 Pro (60.3) and GPT-4.1 (50.4); all seven evaluators ranked DT-RAG first (Kendall's W = 0.97).
Conclusions:
DT-RAG, a curated retrieval-augmented configuration, outperformed frontier-tier general-purpose LLMs on this dental traumatology benchmark, with no confabulated rationales observed among the responses assessed. This approach may be applicable to other well-defined clinical domains with authoritative guidelines.
Clinical Significance:
Curated retrieval augmentation produced accurate, source-traceable and reproducible responses in dental traumatology under benchmark and simulated-scenario conditions, making every error auditable against its source. Clinical safety requires prospective evaluation.

