Related Experiment Videos
Assessment of approximate string matching in a biomedical text retrieval problem.
1Department of Computational Science, National University of Singapore, Blk SOC1, Level 7, 3 Science Drive 2, Singapore 117543, Singapore.
Computers in Biology and Medicine
|August 30, 2005
Summary
The Smith-Waterman algorithm improves biomedical text retrieval by accurately matching medicinal herb names. This method enhances data mining accuracy, achieving high recall and precision.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Text Mining
Background:
- Text-based search is crucial for biomedical data mining and knowledge discovery.
- Character errors in literature can significantly impact the accuracy of data mining processes.
- Developing robust methods to address these errors is essential for reliable biomedical research.
Purpose of the Study:
- To evaluate the effectiveness of the Smith-Waterman algorithm with affine gap penalty for biomedical literature retrieval.
- To assess the algorithm's performance in matching medicinal herb names between herbal medicine and medicinal chemistry literature.
- To determine the optimal string identity level for accurate name matching.
Main Methods:
- The Smith-Waterman algorithm with affine gap penalty was employed for sequence alignment.
- Names of medicinal herbs were extracted from herbal medicine and medicinal chemistry literature.
- The algorithm was tested at various string identity levels ranging from 80% to 100%.
Main Results:
- The Smith-Waterman algorithm demonstrated utility in biomedical text retrieval.
- Optimal performance was achieved at an 88% string identity level.
- At 88% string identity, the algorithm yielded a recall of 96.9% and a precision of 97.3%.
Conclusions:
- The Smith-Waterman algorithm is a valuable tool for enhancing the accuracy of biomedical text retrieval.
- This approach can effectively improve the success rate of data mining in biomedical literature.
- The findings support the use of sequence alignment algorithms for overcoming character errors in scientific texts.