Related Experiment Video
Updated: Jan 17, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Large language models accurately identify immunosuppression in intensive care unit patients
Vijeeth Guggilla1, Mengjia Kang2, Melissa J Bak3
1Institute for Artificial Intelligence in Medicine, Northwestern University Feinberg School of Medicine, Chicago, IL 60611, United States.
Journal of the American Medical Informatics Association : JAMIA
|September 22, 2025
Summary
Large language models (LLMs) show superior accuracy in identifying immunosuppression from clinical notes compared to traditional methods. This advancement improves patient cohort identification for research.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Research
Background:
- Traditional methods like rule-based algorithms and natural language processing (NLP) struggle with accuracy and generalizability in identifying immunosuppression from clinical notes.
- Identifying patients with immunosuppression is crucial for research and clinical management.
Purpose of the Study:
- To compare the performance of large language models (LLMs) against structured data algorithms and NLP approaches for identifying immunosuppression from unstructured clinical notes.
- To evaluate the effectiveness of LLMs, specifically GPT-4o, in identifying immunosuppressive conditions and medications.
Main Methods:
- Utilized hospital admission notes from two cohorts (N=827 and N=200) from different medical centers.
- Evaluated structured data algorithms, NLP approaches, and LLMs (GPT-4o) for identifying 7 immunosuppressive conditions and 6 immunosuppressive medications.
Main Results:
- LLMs, particularly GPT-4o, outperformed structured data algorithms and NLP approaches across all evaluated immunosuppressive conditions and medications.
- GPT-4o achieved high F1 scores (0.51-1) in the primary cohort and demonstrated strong performance in the validation cohort (F1=1 for 8/13 variables).
- Structured data algorithms had variable performance (F1 0.30-0.97), while NLP approaches showed a wide range (F1 0-1).
Conclusions:
- Large language models (LLMs), exemplified by GPT-4o, offer a significant improvement over existing methods for identifying immunosuppression from clinical text.
- LLMs provide a robust and validated tool for enhancing patient cohort identification in clinical research.
