Related Experiment Video
Updated: Jan 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Improving Large Language Models for Clinical Named Entity Recognition via Prompt Engineering.
Arxiv
|October 1, 2025
Summary
This study shows that task-specific prompts significantly improve GPT-3.5 and GPT-4 performance on clinical named entity recognition (NER) tasks, making them more feasible for healthcare applications.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Medicine
- Clinical Informatics
Background:
- Large language models like GPT-3.5 and GPT-4 show promise for clinical natural language processing tasks.
- However, their direct application to clinical named entity recognition (NER) may not achieve optimal performance without specific tuning.
- Evaluating and enhancing GPT models for specialized clinical NER is crucial for their effective integration into healthcare.
Purpose of the Study:
- To quantify the capabilities of GPT-3.5 and GPT-4 on clinical NER tasks.
- To develop and evaluate a task-specific prompt framework to improve GPT model performance in clinical NER.
- To compare the performance of enhanced GPT models against a specialized clinical NER model, BioClinicalBERT.
Main Methods:
- GPT-3.5 and GPT-4 were evaluated on two clinical NER tasks: extracting concepts from clinical notes (MTSamples) and identifying adverse events from safety reports (VAERS).
- A prompt framework was developed, including baseline prompts, guideline-based prompts, error analysis instructions, and few-shot learning samples.
- Model performance was assessed using F1 scores and compared to BioClinicalBERT.
Main Results:
- Without specialized prompts, GPT-3.5 and GPT-4 achieved lower F1 scores on both datasets compared to BioClinicalBERT.
- The proposed prompt framework significantly improved GPT model performance, with GPT-4 achieving F1 scores of 0.861 on MTSamples and 0.736 on VAERS when using all prompt components.
- Despite improvements, the enhanced GPT models still lagged behind BioClinicalBERT but required fewer training samples.
Conclusions:
- Direct application of GPT models to clinical NER is suboptimal.
- A task-specific prompt framework incorporating medical knowledge and training samples substantially enhances GPT models' performance for clinical NER.
- These findings highlight the potential of prompt engineering to adapt large language models for practical clinical applications.
Related Concept Videos
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
