Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Retrieval-Augmented Large Language Models on External Cervical Resorption: A Comparative Study of Gemini
Marc Garcia-Font1, Nicolás Dufey-Portilla2, Fernando Durán-Sindreu1
1Department of Endodontics, Universitat International de Catalunya, School of Dentistry, Barcelona, Spain.
Two AI models, Google Gemini and NotebookLM, showed high accuracy and consistency in answering clinical questions about external cervical resorption. NotebookLM performed slightly better, but retrieval augmentation did not significantly improve responses for these tasks.
Area of Science:
- Artificial Intelligence in Dentistry
- Clinical Decision Support Systems
- Natural Language Processing in Healthcare
Background:
- This study assessed the accuracy and consistency of two Alphabet Inc. large language models: Google Gemini (GG) and NotebookLM (NLM).
- The evaluation focused on answering clinical questions related to external cervical resorption using a retrieval-augmented framework.
- NotebookLM is a document-grounded configuration, while Google Gemini was used in its base configuration.
Purpose of the Study:
- To evaluate the accuracy and consistency of Google Gemini and NotebookLM in responding to clinical questions about external cervical resorption.
- To compare the performance of a base large language model against a document-grounded configuration.
- To determine if retrieval augmentation significantly impacts the quality of responses for structured clinical tasks.
Main Methods:
- Forty-six dichotomous clinical questions on external cervical resorption were created by three endodontic experts.
- Each question was posed to Google Gemini and NotebookLM via three independent user accounts, generating 276 total responses.
- Responses were independently assessed by three endodontic experts against gold standard answers for accuracy and consistency.
Main Results:
- Google Gemini achieved 89% accuracy and 93% consistency.
- NotebookLM achieved 96% accuracy and 90% consistency.
- No statistically significant differences were found between the two models regarding accuracy and consistency.
Conclusions:
- Both NotebookLM and Google Gemini demonstrated high accuracy and consistency in answering clinical questions.
- NotebookLM exhibited a slightly superior performance compared to Google Gemini.
- Retrieval augmentation did not yield significant improvements for these specific structured clinical questions.
More Related Videos
07:22Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy