Related Experiment Videos
Automatic assignment of biomedical categories: toward a generic approach.
1University Hospitals of Geneva, Medical Informatics Service CH-1201, Geneva. Patrick.Ruch@sim.hcuge.ch
Bioinformatics (Oxford, England)
|November 17, 2005
Summary
A new data-independent text categorization system was developed for biomedical texts. This system achieves high accuracy for Medical Subject Headings (MeSH) but lower accuracy for Gene Ontology (GO) categorization.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Information Retrieval
Background:
- Developed a generic text categorization system for automatic biomedical category assignment.
- This system is largely data-independent, unlike traditional data-intensive models.
Purpose of the Study:
- To evaluate the robustness of a novel, data-independent text categorization system.
- To assess the system's performance on two distinct biomedical terminologies: MeSH and GO.
Main Methods:
- Employed a lightweight categorizer with two ranking modules: a pattern matcher and a vector space retrieval engine.
- Utilized both stems and linguistically-motivated indexing units for text analysis.
- Tested the system's effectiveness on Medical Subject Headings (MeSH) and Gene Ontology (GO) datasets.
Main Results:
- Phrase indexing proved effective for both Gene Ontology (GO) and Medical Subject Headings (MeSH) categorization.
- Categorization performance varied significantly based on the controlled vocabulary used.
- Achieved precision above 90% for MeSH categorization at high ranks.
Conclusions:
- The developed system establishes a new baseline for retrieval-based categorizers.
- Categorization accuracy is dependent on the characteristics of the controlled vocabulary.
- Demonstrated high precision for MeSH categorization, while GO categorization showed precision below 20%.