Related Experiment Video
Updated: May 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for deductive qualitative content analysis in dementia-focused embedded pragmatic clinical
Jeffrey Turner1, Spencer Phillips Hey2, Zachary G Baker3
1Center for Gerontology and Healthcare Research, Brown University, 121 South Main Street, Suite 6, Providence, 02903, RI, USA. jeffrey_turner@brown.edu.
Large language models (LLMs) automate thematic coding for intervention implementation studies, significantly reducing time and cost. This AI approach accelerates evaluations in embedded pragmatic clinical trials (ePCTs) while maintaining high accuracy.
Area of Science:
- Implementation Science
- Artificial Intelligence
- Dementia Care Research
Background:
- Thematic coding is crucial for characterizing intervention implementation in embedded pragmatic clinical trials (ePCTs), especially for dementia interventions.
- Manual coding is labor-intensive, hindering timely evaluations.
- Large language models (LLMs) offer potential for automating qualitative data analysis in implementation science.
Purpose of the Study:
- To develop and evaluate an automated algorithm using LLMs (Chat GPT-4o and GPT-4o-mini) for thematic coding of interview transcripts.
- To assess the performance of LLMs compared to human coders in identifying implementation determinants, barriers, and facilitators.
- To quantify the impact of LLM-powered workflows on efficiency (time and cost) in implementation research.
Main Methods:
- A Python-based system was developed to process semi-structured interview transcripts using LLMs.
- The system matched transcript excerpts to an established codebook, undergoing multiple refinement iterations with expert review.
- Performance was compared against human coding, evaluating accuracy, code type matching, and missed excerpts.
Main Results:
- The LLM consistently coded more excerpts than human coders, achieving up to 72.6% matching rates on individual transcripts.
- GPT-4o demonstrated superior performance over GPT-4o-mini, with higher matching rates for descriptive codes (89% vs. 69%).
- The LLM workflow resulted in a 97% reduction in time and a 99% reduction in cost per transcript.
Conclusions:
- LLM-powered thematic analysis aligns well with human coding, offering a reliable supplementary tool for implementation science.
- While human oversight is necessary due to error rates, LLMs significantly enhance efficiency and scalability in qualitative research.
- This automated approach streamlines the analysis of implementation processes, potentially accelerating research and minimizing resource allocation.
Related Concept Videos
Qualitative Analysis
There are two main approaches to qualitative analysis:...
Qualitative Analysis
For instance, group IV...
Dementia
The progression of dementia is generally gradual.
Dementia l: Introduction
