Related Experiment Video
Updated: May 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Source-Based Large Language Models for Preclinical Dermatology Education: Comparative Study
Frank Je-Min Lin1, Sunghun Cho2
1F. Edward Hébert School of Medicine, Uniformed Services University of the Health Sciences, 4301 Jones Bridge Road, Bethesda, US.
Background:
Large language models (LLMs) have gained increasing popularity in medical education, with evidence supporting their educational value when framed through the lens of Cognitive Load Theory (CLT). Source-based LLMs, which explicitly ground responses in user-uploaded material via retrieval-augmented generation (RAG) algorithms, may offer additional educational value by using student-developed materials to conceptualize new areas of learning in a familiar framework. This has applications for areas like medical education in dermatology, which could benefit from inclusive sources and enhanced education to alleviate healthcare gaps. However, no prior studies have examined whether the inclusion of student-authored notes alters the response characteristics of a source-based LLM when responding to medical questions.
Objective:
To conduct an observational, comparative performance evaluation study assessing the accuracy, response reproducibility, and intermodel response similarity of freely available LLMs on text-only Step 1 dermatology questions, and to explore whether providing extensive student-generated notes to a source-based LLM alters these performance characteristics.
Methods:
In December 2024, four LLMs were evaluated: NotebookLM (NLM) with uploaded pre-clerkship study guides (NLM w/ Notes), NLM with an uploaded blank sheet of paper (NLM w/o Notes), ChatGPT-4o mini, and Google Gemini 1.5 Flash. Each model completed three trials of 121 text-based USMLE Step 1 dermatology questions from the AMBOSS question bank. They were evaluated for overall majority-consensus accuracy, accuracy by question difficulty, intertrial reproducibility, and agreement in answer choice selection between models. Differences were analyzed through a Cochran's Q omnibus test and subsequent pairwise McNemar tests with Benjamini-Hochberg correction. Response reproducibility and intermodel agreement were analyzed through Fleiss's Kappa statistics with 95% confidence intervals.
Results:
ChatGPT-4o mini achieved the highest overall majority-consensus accuracy (84.3%). NLM w/ Notes demonstrated the highest intertrial reproducibility (Fleiss's κ = 0.927, 95% CI [0.875-0.978]) and strong performance on lower-difficulty questions, but comparatively reduced accuracy on higher-difficulty items. NLM w/o Notes exhibited significantly higher omission rates (10.5% vs ≤1.65% for other models) than other tested LLMs. Sensitivity analysis excluding omissions increased NLM w/o Notes' accuracy from 66.9% to 77.8%, matching NLM w/ Notes' performance. Intermodel agreement was significantly higher between NLM w/ Notes and ChatGPT-4o mini compared to NLM w/o Notes and Gemini 1.5 Flash.
Conclusions:
Provision of student-generated notes substantially increased response reproducibility in a source-based LLM, likely reflecting consistent retrieval of similar source excerpts across trials. However, note-grounding appeared to constrain performance on higher-difficulty questions, suggesting a RAG algorithm retrieval error when question stems excluded characteristic 'key words' present in lower-difficulty items. The results highlight potential challenges of a student-level, CLT-grounded educational LLM that must deal with non-expert-curated notes, balance source utilization and internal reasoning, and meaningfully appraise uploaded sources to assess a student's individual learning gaps.

