Related Experiment Video
Updated: Nov 12, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Building a semantically annotated corpus for chronic disease complications using two document types
1Department of Computer Science and Engineering, Royal Commission for Jubail and Yanbu, Yanbu University College, Yanbu Industrial City, Saudi Arabia.
We created the PrevComp corpus, a unique dataset from electronic health records and Twitter, to identify disease complications and risk factors for hypertension and diabetes prevention. This resource aids text-mining tool development for better health insights.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Public Health
Background:
- Electronic health records (EHRs) and social media (Twitter) contain valuable patient health information, including disease complications and risk factors.
- Understanding disease risk factors is crucial for prevention and managing complications, particularly for conditions like hypertension and diabetes.
- Developing effective text-mining tools for health information extraction requires high-quality annotated data.
Purpose of the Study:
- To develop the PrevComp corpus, an annotated dataset for identifying disease complications, risk factors, and prevention measures.
- To focus on the interaction between hypertension and diabetes within the biomedical domain.
- To facilitate the creation of advanced text-mining tools for extracting critical health insights.
Main Methods:
- Integrated narrative text from electronic health records (EHRs) and tweets from Twitter.
- Developed a unique annotation scheme for disease complications, risk factors, and prevention measures, guided by domain experts.
- Conducted expert-driven annotation to ensure high-quality data, achieving F-scores of 0.60 for EHRs and 0.75 for tweets.
Main Results:
- Successfully created the PrevComp corpus, a novel resource combining EHR and Twitter data specific to hypertension and diabetes.
- The corpus is annotated for disease complications, risk factors, and prevention strategies.
- Achieved high inter-annotator agreement, indicating reliable and high-quality annotations.
Conclusions:
- The PrevComp corpus is a valuable, unique resource for advancing text-mining in the biomedical domain, specifically for hypertension and diabetes research.
- This annotated corpus will support the development of tools to extract vital information for disease risk factor management and prevention.
- The integration of EHR and Twitter data offers a comprehensive approach to capturing diverse health-related narratives.
More Related Videos
Related Concept Videos
Chronic Obstructive Pulmonary Disease
Smoking is a primary risk factor for COPD, with over 80% of patients having a history of it. Patients typically experience progressive dyspnea or labored breathing, frequent coughing, and recurrent pulmonary infections. Many eventually succumb to respiratory failure, characterized by...
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation II: POMR
Chronic Pancreatitis II: Collaborative Care
Assessment:
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...

