Extracting COVID-19 Diagnoses and Symptoms From Clinical Text: A New Annotated Corpus and Neural Event Extraction

Arxiv
|December 10, 2020
PubMed

Insights

This study introduces a new corpus and model for extracting COVID-19 information from clinical notes. Automatically extracted symptoms improve prediction of test results.

Area of Science:

  • Medical Informatics
  • Natural Language Processing
  • Epidemiology

Background:

  • The COVID-19 pandemic necessitates advanced methods for analyzing clinical data.
  • Free-text clinical notes are rich in information but challenging to process at scale.
  • Existing methods lack the ability to efficiently extract comprehensive COVID-19 related data from unstructured text.

Approach:

  • Developed the COVID-19 Annotated Clinical Text (CACT) Corpus with 1,472 notes.
  • Introduced a span-based event extraction model for joint information extraction.
  • Achieved high F1 scores (0.83-0.97) for identifying COVID-19 events and symptoms.

Key Points:

  • The CACT Corpus provides detailed annotations for COVID-19 diagnoses, testing, and clinical presentation.
  • The event extraction model demonstrates strong performance in identifying key clinical phenomena.
  • Automatically extracted symptom data enhances the prediction of COVID-19 test results.

Conclusions:

  • The CACT Corpus and associated model facilitate large-scale analysis of COVID-19 clinical text.
  • Automated information extraction is crucial for understanding and managing the pandemic.
  • Integrating extracted symptoms with structured data improves predictive accuracy for COVID-19 outcomes.