Related Experiment Video
Updated: Jul 4, 2026

14:34
A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Evaluating Prompt Strategies for LLM-Based De-Identification of German Discharge Letters: A Feasibility Study Using
Florin Dominik Teschner1, Hung Manh Nguyen1, Martin Sedlmayr1
1Institute for Medical Informatics and Biometry, Faculty of Medicine Hospital Carl Gustav Carus, TUD Dresden University of Technology, Dresden, Germany.
Studies in Health Technology and Informatics
|July 3, 2026
Summary
Optimizing prompts for large language models (LLMs) significantly improves clinical text de-identification, achieving high F1-scores. Careful prompt design is crucial for effective LLM-based de-identification in healthcare.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Healthcare
- Medical Informatics
Background:
- Clinical text de-identification is vital for secondary data use.
- Heterogeneous clinical documentation presents significant de-identification challenges.
- Large Language Models (LLMs) offer potential solutions for automated de-identification.
Purpose of the Study:
- To evaluate LLM-based de-identification of synthetic German discharge letters.
- To compare the impact of different prompting strategies on de-identification performance.
- To assess the effectiveness of GPT-4o and GPT-OSS models for this task.
Main Methods:
- Utilized synthetic German discharge letters (GraSCCo) for evaluation.
- Employed four distinct prompting strategies with GPT-4o.
- Included a comparative analysis with a GPT-OSS model.
- Assessed performance using precision, recall, F1-score, false positives, and text reduction.
Main Results:
- Baseline LLM setup resulted in excessive text loss.
- Optimized prompting with GPT-4o achieved a high F1-score of 0.93.
- GPT-OSS also demonstrated strong performance with an F1-score of 0.90.
- Prompt refinement showed a greater impact than structural preprocessing.
Conclusions:
- LLM-based de-identification is feasible for clinical text.
- Careful prompt engineering is critical for achieving high de-identification quality.
- Validation on real-world clinical data is necessary for practical implementation.
