Related Experiment Video
Updated: Aug 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Model-Assisted Thematic Coding in Medical Education Research: Comparative Methodological Study
Alexa DeRegnaucourt1, Katherine Miller Jennings1, Andrew Zahn1
1College of Medicine, University of Cincinnati, Cincinnati, OH, United States.
Background:
While large language model (LLM)-assisted qualitative analysis could improve the efficiency and scalability of feedback-driven curricular refinement in medical education, how best to leverage LLMs for qualitative analysis while ensuring quality outputs remains an open question. Prior work has demonstrated the feasibility of using LLMs for inductive and deductive coding tasks, but more needs to be known about how LLM-assisted thematic coding can best be deployed in a medical education context to maximize its strengths and guard against its weaknesses.
Objective:
Our study evaluated LLM performance in inductive code generation and in the deductive application of a human codebook, using a student focus-group transcript, to propose a model for AI collaboration in qualitative analysis.
Methods:
The qualitative data for this study consisted of a 1-hour focus group with 4 second-year medical students discussing a required AI-driven clinical-scenario tool (2-Sigma). Three human coders conducted an inductive thematic analysis. Using the same transcript, GPT-4o (version gpt-4o-2024-11-20; OpenAI) generated inductive codes and applied the human codebook deductively. The researchers compared the alignment between the AI inductive codes and the human consensus codebook using 3 categories: agreement, reasonable alternative, and not reasonable. Interrater reliability of AI deductive coding was evaluated using percent agreement and Cohen κ, with textual audits of discrepancies, including "misses" (failed to apply appropriate codes) and "misfires" (inappropriately applied codes). Analysis took place between February and July 2025.
Results:
In the inductive condition, GPT-4o generated 137 initial codes, of which 31.4% (n=43) demonstrated agreement with human codes, 26.3% (n=36) represented reasonable alternatives, and 42.3% (n=58) were classified as not reasonable. In the deductive condition, mean percent agreement for AI application of human codes was 96% (SD 4%, range 79%-100%) and the mean κ was 0.71 (SD 0.26, range 0-1.00). Of all 2352 coding decisions, there were 57 (2.4%) misfires and 28 (1.2%) misses; common patterns included overinterpretation of tone, failure to recognize continued ideas across excerpts, and difficulty distinguishing hypothetical vs experienced features. Based on our findings, we suggest a roadmap that retains human interpretive control while leveraging AI scalability: humans first develop a contextually grounded codebook through inductive analysis, then use AI both as a creative partner to surface alternative codes and as a tool to apply the validated codebook across the dataset.
Conclusions:
With targeted human oversight, an LLM reliably applied an existing codebook and generated additional inductive codes. These findings support a proposed workflow in which AI serves as an additional perspective within human-driven qualitative analysis, offering a scalable adjunct for qualitative analysis in medical education. Validation across larger and more diverse datasets will help confirm the generalizability of this approach.