Related Experiment Video
Updated: Apr 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using a Large Language Model to Extract Information from Student Submitted Free-Text Feedback
Nikola Košćica1, Colleen Gillespie1, Tyler Webster1
1NYU Grossman School of Medicine, New York, NY USA.
None:
Student feedback is essential to curriculum evaluation. While methods for analyzing quantitative feedback data are readily available and easy to implement, methods for analyzing text-based, qualitative feedback data are less widely available, requiring more time, effort, and expertise. And yet, students' responses to open-ended questions hold great value for curriculum refinement because narrative comments can identify un-anticipated areas of concern that closed-ended rating scales might miss and often provide specific suggestions for improvement. In this paper, we describe efforts to analyze the feasibility and accuracy of using a Large Language Model (ChatGPT 4o) to analyze medical student comments in response to a question asking them to identify basic science topics they found challenging. ChatGPT 4o was used to categorize and summarize students' identification of and explanations for these challenging topics. We describe the specific prompts used to generate and refine results and then conducted a series of experiments to explore consistency, accuracy, and meaningfulness: (1) reviewing the consistency of 10 replications of the ChatGPT 4o request; (2) comparing "expert" human ratings of topic categories with ChatGPT's categorization; and (3) comparing "expert" human analyses of the explanations for a challenging topic with those generated by ChatGPT. Overall, we found the LLM output to be useful, fairly closely aligned with human experts, and easy to implement. However, results were not perfectly replicated across multiple trials and we found some differences between human and LLM analyses. Our use case is well suited to the current capabilities of genAI models in that summaries can be rapidly and easily generated with sufficient (but not perfect) consistency and accuracy to support continuous quality improvement of basic science curriculum.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s40670-025-02615-1.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Comparing Experimental Results: Student's t-Test
Student t Distribution
The Student t distribution was developed by William S. Goset (1876–1937) of the...