Related Experiment Videos
Combining NLP and probabilistic categorisation for document and term selection for Swiss-Prot medical annotation
Pavel B Dobrokhotov1, Cyril Goutte, Anne-Lise Veuthey
1Swiss Institute of Bioinformatics, CMU, 1 Michel-Servet - CH-1211 Geneva 4, Switzerland. Pavel.Dobrokhotov@isb-sib.ch
Bioinformatics (Oxford, England)
|July 12, 2003
Summary
This study introduces a Natural Language Processing (NLP) and probabilistic classification method to improve document relevance for Swiss-Prot annotation. The approach enhances database curation by identifying key terms and re-ranking search results.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing
Background:
- Manual database annotation is time-consuming.
- Efficiently searching and ranking relevant publications is crucial for biological databases.
Purpose of the Study:
- To develop an automated method for re-ranking PubMed documents for Swiss-Prot annotation.
- To identify significant terms within documents relevant to Swiss-Prot annotation using NLP and probabilistic classification.
Main Methods:
- Applied Natural Language Processing (NLP) and probabilistic classification.
- Utilized a Probabilistic Latent Categoriser (PLC) for document re-ranking.
- Employed Kullback-Leibler symmetric divergence to identify discriminating terms.
Main Results:
- Achieved 69% recall and 59% precision for relevant documents in a representative query.
- Identified key terms that contribute to document relevance for Swiss-Prot annotation.
- Demonstrated the value of term contribution analysis for understanding classification results.
Conclusions:
- The developed method effectively re-ranks documents and identifies significant terms for Swiss-Prot annotation.
- The approach aids curators in understanding classification outcomes and refining linguistic pre-processing.
- This enhances the overall performance and efficiency of biological database annotation.