Related Experiment Video
Updated: Sep 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Improving Drug Identification in Overdose Death Surveillance using Large Language Models.
Natural language processing (NLP) models, specifically BioClinicalBERT, accurately classify drug-involved overdose deaths from free-text reports. This enhances timely surveillance, overcoming limitations of manual coding for emerging substance use trends.
Area of Science:
- Public Health Surveillance
- Computational Linguistics
- Data Science
Background:
- Drug-related deaths, primarily fentanyl-driven, necessitate rapid and precise surveillance in the U.S.
- Current methods relying on manual coding of coroner reports into ICD-10 classifications cause data delays and loss.
- Existing natural language processing (NLP) applications for overdose surveillance have shown limitations.
Purpose of the Study:
- To evaluate and compare various NLP models for classifying specific drug involvement from unstructured death certificate text.
- To assess the performance of traditional machine learning, general-domain BERT, and large language models (LLMs) against fine-tuned clinical NLP models.
- To determine the efficacy of NLP in automating and enhancing overdose surveillance data extraction.
Main Methods:
- Utilized a dataset of 35,433 U.S. death records from 2020 for training and internal testing.
- Performed external validation on a separate dataset of 3,335 records from 2023-2024.
- Compared traditional classifiers, Bidirectional Encoder Representations from Transformers (BERT), BioClinicalBERT, Qwen 3, and Llama 3 using macro-averaged F1 scores.
Main Results:
- Fine-tuned BioClinicalBERT models achieved near-perfect performance (macro F1 >=0.998) on the internal test set.
- External validation demonstrated the robustness of BioClinicalBERT (macro F1=0.966), outperforming other evaluated models.
- NLP models significantly outperformed conventional machine learning and general-domain BERT and LLMs.
Conclusions:
- Fine-tuned clinical NLP models, like BioClinicalBERT, provide a highly accurate and scalable solution for classifying drug-involved overdose deaths from free-text reports.
- These NLP methods can substantially accelerate surveillance workflows, surpassing the limitations of manual ICD-10 coding.
- The approach supports near real-time detection of emerging substance use trends, improving public health response.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
10:17High-throughput and Comprehensive Drug Surveillance Using Multisegment Injection-Capillary Electrophoresis-Mass Spectrometry
Published on: April 23, 2019
Related Concept Videos
Drug Discovery: Overview
Enhanced Elimination of Poison
Antidotes serve a crucial role in counteracting the effects of poison by inhibiting enzymes responsible for producing harmful drug metabolites. In some cases, these toxic metabolites can be neutralized by endogenous cosubstrates, which are maintained at specific concentrations to prevent interaction with cellular macromolecules and subsequent cell death.
Renal excretion is the...
Drug Nomenclature
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Analysis of Population Pharmacokinetic Data