Related Experiment Video
Updated: Feb 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Enhanced Large Language Models on Prosthodontic Multiple-Choice Questions
Shenghan Gao1, Zi-Ang Wang2, Zihan Gao1
1Department of Prosthodontics, Peking University School and Hospital of Stomatology & National Center for Stomatology & National Clinical Research Center for Oral Diseases & National Engineering Research Center of Oral Biomaterials and Digital Medical Devices & Beijing Key Laboratory of Digital Stomatology & NHC Key Laboratory of Digital Stomatology & NMPA Key Laboratory for Dental Materials, Beijing, PR China.
Enhanced large language models (LLMs) using retrieval-augmented generation (RAG), in-context learning (ICL), and majority voting showed improved accuracy on Chinese prosthodontic questions but not English ones. These enhanced LLMs show potential for dental education tasks.
Area of Science:
- Artificial Intelligence in Dentistry
- Natural Language Processing for Dental Education
- Large Language Model Performance Evaluation
Background:
- Large language models (LLMs) are increasingly explored for specialized domains like dentistry.
- Enhancement strategies such as retrieval-augmented generation (RAG), in-context learning (ICL), and majority voting aim to improve LLM accuracy and reliability.
- Evaluating LLM performance on domain-specific questions is crucial for their practical application.
Purpose of the Study:
- To assess the effectiveness of retrieval-augmented generation (RAG), in-context learning (ICL), and majority voting in enhancing large language models (LLMs) for prosthodontic questions.
- To compare the performance of enhanced LLMs against their base versions using both Chinese and English prosthodontic question sets.
- To analyze error types and identify areas where LLM enhancements yield significant improvements.
Main Methods:
- Two base large language models (OpenAI o1, DeepSeek-R1) were enhanced using RAG, ICL, and majority voting techniques.
- Standardized Chinese-language multiple-choice questions (C-MCQs) and English-language multiple-choice questions (E-MCQs) in prosthodontics were used for evaluation.
- Performance was measured by answer correctness, error analysis, and statistical comparison (χ² tests) between base and enhanced models.
Main Results:
- Enhanced LLMs demonstrated statistically significant accuracy improvements on C-MCQs (P < .001) and a reduction in knowledge-based errors (P < .008).
- While enhanced LLMs showed higher accuracy on E-MCQs, the difference was not statistically significant (P = .145).
- The benefits of LLM enhancement, particularly in accuracy and error reduction, were primarily observed in the Chinese-language question set.
Conclusions:
- Inference-time enhancement strategies (RAG, ICL, majority voting) significantly boost LLM accuracy for Chinese prosthodontic questions.
- These enhancements effectively mitigate certain errors but do not consistently yield statistically significant improvements for English questions.
- Enhanced LLMs show promise for processing dental-related tasks and hold potential for applications in dentistry education.
More Related Videos
07:32Author Spotlight: 3D Movement Assessment of Maxillary Posterior Teeth in Clear Aligner Treatment
Published on: February 23, 2024
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024