Related Experiment Video
Updated: Jun 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human
Michael S Deiner1, Vlad Honcharov2,3, Jiawei Li4
1Department of Ophthalmology and Francis I Proctor Foundation, University of California San Francisco, San Francisco, CA, United States.
Large language models (LLMs) can efficiently analyze social media health content, identifying relevant topics and reasonable themes. While not fully replicating human depth, LLMs show promise for public health social listening.
Area of Science:
- Computational linguistics
- Public health informatics
- Social media analytics
Background:
- Manual analysis of social media provides public health insights but is time-intensive.
- Generative large language models (LLMs) offer potential for summarizing and interpreting large text volumes.
- The effectiveness of LLMs in discerning subtle health meanings from social media remains unclear.
Purpose of the Study:
- To assess if LLMs can perform topic model selection and thematic analysis on social media content.
- To compare LLM performance against human subject matter experts from a prior study.
- To determine if LLMs can reasonably identify health-related themes in social media data.
Main Methods:
- Replicated a prior study's research question and social media content using three LLMs (GPT4-32K, Claude-instant-100K, Claude-2-100K).
- Compared LLM topic selection and theme identification with manual human analyses.
- Evaluated LLM consistency and inter-model agreement.
Main Results:
- LLMs ranked previously human-identified topics highly, exceeding chance levels (P<.001).
- LLMs identified relevant themes with low hallucination rates, deemed reasonable by experts.
- Variability in theme identification was observed between different LLMs and repeated analyses.
Conclusions:
- LLMs efficiently process large social media health datasets and extract reasonable themes.
- LLMs show significant potential for automated public health social listening.
- Further validation is needed to match the depth of human expert analysis in theme extraction.
More Related Videos
Related Concept Videos
Qualitative Analysis
For instance, group IV...
Stereotype Content Model
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Longitudinal Research
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...

