Related Experiment Video
Updated: May 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models Accurately Identify People Who Inject Drugs From Infectious Diseases Discharge Summaries in an
David Goodman-Meza1,2, Marianne Martinello2,3, Jeffrey Masters2,4
1St Vincent's Hospital, Sydney, Australia.
Drug and Alcohol Review
|May 14, 2026
Summary
Large language models (LLMs) effectively identify people who inject drugs (PWID) from clinical notes, significantly outperforming traditional ICD codes. This advancement aids in better public health surveillance and targeted interventions for PWID.
Area of Science:
- Natural Language Processing in Healthcare
- Clinical Informatics
- Public Health Surveillance
Background:
- People who inject drugs (PWID) are at high risk for infections, but International Classification of Diseases (ICD) codes inadequately identify this population.
- Unstructured clinical text contains valuable data, and large language models (LLMs) offer a novel method for information extraction.
- Evaluating LLM diagnostic performance is crucial for improving identification of at-risk populations in healthcare settings.
Purpose of the Study:
- To assess the diagnostic accuracy of various off-the-shelf large language models (LLMs) in identifying people who inject drugs (PWID).
- To compare the performance of LLMs against traditional International Classification of Diseases (ICD) codes for PWID identification.
- To evaluate the ability of LLMs to extract specific drug use information and treatment data from hospital discharge summaries.
Main Methods:
- A cross-sectional study analyzed de-identified hospital discharge summaries from an Infectious Diseases service (2018-2022).
- Manual annotation was performed by a single reviewer to identify PWID status, drugs used, injection recency, and opioid agonist therapy.
- Eight LLMs were evaluated using prevalence-weighted average-F1 scores, with diagnostic metrics and bootstrapped 95% confidence intervals calculated.
Main Results:
- Manual review identified 17.1% of 859 first admissions as PWID; ICD codes demonstrated low sensitivity (≤0.32) but high specificity (≥0.97).
- The top-performing LLM, Llama 3.3, achieved a prevalence-weighted average-F1 score of 0.845 for PWID identification, with high sensitivity (0.819) and specificity (0.999) for injecting drug use.
- LLMs achieved near-perfect identification (F1 > 0.973) for heroin, methamphetamine, cannabis, and methadone, but showed lower accuracy for illicit prescription opioids (F1=0.400) and benzodiazepines (F1=0.606).
Conclusions:
- Large language models demonstrate high accuracy in identifying people who inject drugs (PWID) from clinical discharge summaries, surpassing the performance of ICD codes.
- While effective for many substances, LLMs require further task-specific tuning and external validation for accurate identification of certain drug classes.
- Integrating LLM insights with structured data holds promise for enhancing public health surveillance and developing targeted interventions for PWID.