Related Experiment Video
Updated: Jan 24, 2026

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Improving chemical disease relation extraction with rich features and weakly labeled data
Yifan Peng1, Chih-Hsuan Wei2, Zhiyong Lu2
1National Center for Biotechnology Information, Bethesda, MD 20894 USA ; Computer and Information Sciences, University of Delaware, Newark, DE 19716 USA.
This study introduces an advanced text-mining system for identifying chemical-induced disease (CID) relations, improving drug discovery and chemical safety. The system achieves state-of-the-art performance by combining machine learning with existing knowledge bases.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Computational Biology
Background:
- Identifying chemical-disease relationships is crucial for drug discovery and chemical safety.
- Biomedical literature contains vast information on these relationships, necessitating automated extraction methods.
- Existing methods for chemical-induced disease (CID) relation extraction require improvement.
Purpose of the Study:
- To enhance the state-of-the-art in biomedical relation extraction, specifically for CID relations.
- To develop an automatic system for extracting CID relations from biomedical literature.
- To leverage advances in named entity recognition and BioCreative efforts.
Main Methods:
- A Support Vector Machine (SVM) based approach utilizing a rich feature set.
- Incorporation of novel statistical features, linguistic knowledge, and domain resources.
- Integration of a rule-based system's output and automatically generated labeled text from knowledge bases.
Main Results:
- The system achieved an F-score of 57.51% on the BioCreative V dataset, outperforming previous methods.
- Performance improved to 61.01% F-score when augmented with automatically generated weakly labeled data.
- The approach combines the strengths of rule-based and machine learning systems.
Conclusions:
- The developed text-mining approach achieves state-of-the-art performance in disease-chemical relation extraction.
- This work highlights the value of utilizing curated document-level annotations from existing biomedical databases.
- The findings suggest a more effective way to develop text-mining systems by leveraging overlooked resources.
Related Concept Videos
Weak Base Solutions
Weak Acid Solutions
Titration of a Weak Acid with a Weak Base
As a result, there is no simple...
Gastroesophageal Reflux Disease II: Clinical Features and Management
Clinical Manifestations
GERD presents itself in a multitude of ways, with symptoms varying from person to person. The hallmark symptoms are...
Titration Calculations: Weak Acid - Strong Base
For the titration of 25.00 mL of 0.100 M CH3CO2H with 0.100 M NaOH, the reaction can be represented as:
Chemical Formulas

