Related Experiment Video
Updated: Jul 15, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
605
A Natural Language Processing Model for COVID-19 Detection Based on Dutch General Practice Electronic Health Records
Maarten Homburg1, Eline Meijer1,2, Matthijs Berends1,3,4
1Department of Primary- and Long-Term Care, University Medical Center Groningen, Groningen, Netherlands.
Journal of Medical Internet Research
|October 4, 2023
Summary
A BERT model accurately identified COVID-19 consultations in Dutch general practice electronic health records, even predicting cases before official confirmation. This natural language processing approach offers a powerful tool for early disease surveillance and outbreak detection.
Area of Science:
- Computational linguistics
- Medical informatics
- Epidemiology
Background:
- Natural language processing (NLP) models like BERT show potential for improving disease identification from electronic health records (EHRs).
- The COVID-19 pandemic highlighted challenges in disease surveillance, particularly with unstructured EHR data.
- General practitioners' EHRs in the Netherlands contain valuable information but are difficult to analyze due to free-text entries.
Purpose of the Study:
- To develop and validate a BERT model for detecting COVID-19 consultations within Dutch general practice EHRs.
- To assess the model's accuracy and reliability in identifying COVID-19 cases.
Main Methods:
- A BERT model was pre-trained on Dutch language data and fine-tuned using a dataset of COVID-19 and non-COVID-19 general practitioner consultations.
- Model performance was evaluated on an independent test set and validated through external EHR data, PCR test results, and hospitalization data.
- The study utilized 300,359 general practitioner consultations for model development.
Main Results:
- The developed BERT model achieved high accuracy (0.97) and F1-score (0.90) for COVID-19 consultation detection.
- External validation demonstrated comparable high performance, while PCR validation showed high recall but lower precision and specificity.
- The model significantly correlated with COVID-19 hospitalizations and could predict cases weeks before the first confirmed case in the Netherlands.
Conclusions:
- The validated BERT model accurately identifies COVID-19 cases in general practice EHRs, even preceding confirmed diagnoses.
- This demonstrates the potential of NLP models for early disease outbreak detection and effective disease surveillance.
- The study provides a blueprint for using similar models for early recognition of various illnesses.
Keywords:
BERT modelCOVID-19EHRNLPdisease identificationelectronic health recordsmodel developmentmultidisciplinarynatural language processingpredictionprimary carepublic health
