Related Experiment Video
Updated: Jun 13, 2025

Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
Published on: July 7, 2023
Detecting the clinical features of difficult-to-treat depression using synthetic data from large language models
Isabelle Lorge1, Dan W Joyce2, Niall Taylor1
1Department of Psychiatry, University of Oxford, UK.
This study developed a BERT-based model using synthetic data to identify prognostic factors for difficult-to-treat depression (DTD) from electronic health records. The model successfully extracts key DTD predictors from clinical text, showing promise for automated healthcare applications.
Area of Science:
- Computational psychiatry
- Natural Language Processing in Healthcare
- Machine Learning for Clinical Decision Support
Background:
- Difficult-to-treat depression (DTD) presents a significant clinical burden, necessitating improved identification methods.
- Routinely collected electronic health record (EHR) narrative data contains valuable prognostic factors for DTD.
- Current methods for extracting these factors are limited, requiring manual annotation and expert review.
Purpose of the Study:
- To develop and validate a tool for interrogating free-text EHR data to identify prognostic factors of DTD.
- To leverage Large Language Models (LLMs) and span extraction techniques for automated factor detection.
- To assess the feasibility of training a model on synthetic data for clinical data extraction.
Main Methods:
- Utilized GPT3.5 for generating synthetic EHR data.
- Trained a BERT-based span extraction model incorporating a Non-Maximum Suppression (NMS) algorithm.
- Model trained to identify and label positive and negative prognostic factors for DTD.
Main Results:
- Achieved 0.70 F1 score on clinical EHR data for extracting 20 DTD predictors.
- Demonstrated high performance (0.85 F1, 0.95 precision) on key DTD factors including abuse history, family history, illness severity, and suicidality.
- Successfully trained a model exclusively on synthetic data to extract prognostic factors from clinical data.
Conclusions:
- It is feasible to train a machine learning model on synthetic data for extracting clinical prognostic factors from EHRs.
- The developed span extraction model shows significant promise for automated DTD detection and analysis in healthcare.
- This approach could reduce reliance on costly human expert annotations for sensitive medical data.
Related Concept Videos
Antidepressant Drugs: MAOIs and Other Agents
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
Depression: Overview
Long-term Depression
Calcium Ion Concentration Mechanism
If over...
Depressive Disorders: MDD and Dysthymia

