Related Experiment Video
Updated: Nov 26, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Extracting COVID-19 Diagnoses and Symptoms From Clinical Text: A New Annotated Corpus and Neural Event Extraction
Insights
This study introduces a new corpus and model for extracting COVID-19 information from clinical notes. Automatically extracted symptoms improve prediction of test results.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Epidemiology
Background:
- The COVID-19 pandemic necessitates advanced methods for analyzing clinical data.
- Free-text clinical notes are rich in information but challenging to process at scale.
- Existing methods lack the ability to efficiently extract comprehensive COVID-19 related data from unstructured text.
Approach:
- Developed the COVID-19 Annotated Clinical Text (CACT) Corpus with 1,472 notes.
- Introduced a span-based event extraction model for joint information extraction.
- Achieved high F1 scores (0.83-0.97) for identifying COVID-19 events and symptoms.
Key Points:
- The CACT Corpus provides detailed annotations for COVID-19 diagnoses, testing, and clinical presentation.
- The event extraction model demonstrates strong performance in identifying key clinical phenomena.
- Automatically extracted symptom data enhances the prediction of COVID-19 test results.
Conclusions:
- The CACT Corpus and associated model facilitate large-scale analysis of COVID-19 clinical text.
- Automated information extraction is crucial for understanding and managing the pandemic.
- Integrating extracted symptoms with structured data improves predictive accuracy for COVID-19 outcomes.
Abstract:
Coronavirus disease 2019 (COVID-19) is a global pandemic. Although much has been learned about the novel coronavirus since its emergence, there are many open questions related to tracking its spread, describing symptomology, predicting the severity of infection, and forecasting healthcare utilization. Free-text clinical notes contain critical information for resolving these questions. Data-driven, automatic information extraction models are needed to use this text-encoded information in large-scale studies. This work presents a new clinical corpus, referred to as the COVID-19 Annotated Clinical Text (CACT) Corpus, which comprises 1,472 notes with detailed annotations characterizing COVID-19 diagnoses, testing, and clinical presentation. We introduce a span-based event extraction model that jointly extracts all annotated phenomena, achieving high performance in identifying COVID-19 and symptom events with associated assertion values (0.83-0.97 F1 for events and 0.73-0.79 F1 for assertions). In a secondary use application, we explored the prediction of COVID-19 test results using structured patient data (e.g. vital signs and laboratory results) and automatically extracted symptom information. The automatically extracted symptoms improve prediction performance, beyond structured data alone.

