Related Experiment Video
Updated: Oct 3, 2026

Network Analysis of Foramen Ovale Electrode Recordings in Drug-resistant Temporal Lobe Epilepsy Patients
Published on: December 18, 2016
Reproducible synthetic clinical letters for seizure frequency information extraction
Yujian Gan1, Stephen H Barlow2, Ben Holgate2
1Department of Basic and Clinical Neuroscience, Institute of Psychiatry Psychology and Neuroscience, King's College London, 16 De Crespigny Park, London, SE5 8AB, United Kingdom; School of Electronics, Electrical Engineering and Computer Science, Queen's University Belfast, 16A Malone Road, Belfast, BT9 5BN, Northern Ireland, United Kingdom; Centre for Epilepsy, King's College Hospital NHS Foundation Trust, Denmark Hill, London, SE5 9RS, United Kingdom.
Objective:
Seizure-frequency information is critical for epilepsy research and clinical decision-making, yet it is often documented in highly variable free-text neurology clinic letters that are difficult to share and annotate at scale. This study focused on adult epilepsy clinic letters from patients aged 18 years or older. We aimed to develop a reproducible and privacy-preserving framework for seizure-frequency information extraction by training models using fully synthetic, task-faithful clinical letters.
Methods:
We designed a structured label scheme that captures common linguistic patterns of seizure burden, including explicit rates, numeric ranges, cluster-based descriptions, seizure-free durations, unknown frequency, and explicit no-seizure statements. A high-capacity teacher language model was used to generate NHS-style synthetic epilepsy follow-up letters paired with normalized labels, step-by-step rationales, and evidence spans. We fine-tuned multiple open-weight large language models (4-14B parameters) to extract seizure frequency from full letters using either direct numeric outputs or structured labels, optionally augmented with evidence-grounded explanations. Performance was evaluated on a clinician double-checked held-out set of real clinic letters using both fine-grained and pragmatic frequency groupings.
Results:
Models trained purely on synthetic letters generalized to real-world letters, and structured label targets consistently outperformed direct numeric regression. With 15,000 synthetic training letters, representative models achieved micro-F1 up to 0.79 (fine-grained) and 0.85 (pragmatic) on the real test set, while a medically oriented 4B model reached 0.79 and 0.86, respectively. Evidence-grounded outputs supported rapid clinical verification and facilitated error analysis.
Conclusion:
Task-specific synthetic clinical letters with structured, evidence-grounded supervision enabled reproducible development of seizure-frequency extractors while reducing reliance on sensitive adult patient text. These findings support this approach for adult epilepsy clinic-letter extraction, but do not establish generalizability to paediatric notes or other clinical domains.
Related Concept Videos
Seizures ll: Types
Seizures l: Introduction
Epilepsy ll: Types

