Related Experiment Video
Updated: Jul 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models as data-driven engines for benchmarking preventive and clinical knowledge in Chinese dental
Yong Zeng1,2,3, Xinyi Hu2, Wei Liu2
1School of Dentistry, Shenzhen University Medical School, Shenzhen, Guangdong, China.
Introduction:
Digital and data-driven technologies are increasingly shaping preventive oral healthcare and dental education. While prior studies have explored large language model (LLM) performance in non-English examination settings, their role in benchmarking standardized clinical knowledge against real student performance under identical institutional conditions remains insufficiently examined.
Methods:
This study evaluated GPT-4o, GPT-4.5, and DeepSeek-R1 using 300 standardized dental competency items from institutional examinations (2023-2025). Model performance was benchmarked against data from a real student cohort who completed the same examinations, and intra-model consistency was assessed across multiple independent trials.
Results:
DeepSeek-R1 achieved the highest overall accuracy (92.33%), significantly outperforming GPT-4.5 and GPT-4o in 2023 and 2024, while showing comparable performance to GPT-4.5 in 2025. Model type was the predominant factor associated with accuracy differences, whereas knowledge domain did not significantly affect results. DeepSeek-R1 and GPT-4.5 also demonstrated greater response consistency than GPT-4o.
Discussion:
These findings suggest that optimized LLMs demonstrate strong alignment with standardized examination criteria and student performance patterns, supporting their potential role as supplementary tools for curriculum evaluation in preventive dental education. Caution is warranted in generalizing these results beyond the single-institutional, text-only, Mandarin-language context.
