Related Experiment Videos
Effective query filtering for fast homology searching.
1Department of Computer Science, RMIT University, Melbourne, Australia. hugh@cs.rmit.edu.au
Summary
Filtering protein queries can harm homology search accuracy. A new method, cafefilter, masks low complexity regions effectively, improving retrieval effectiveness in large-scale database tests.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Rapid homology searching is crucial in bioinformatics.
- Low complexity regions in query sequences are often masked to improve search accuracy.
- Current filtering techniques are widely applied but their effectiveness can vary.
Purpose of the Study:
- To evaluate the impact of unselective query filtering on homology search accuracy.
- To introduce and assess a novel filtering technique, cafefilter.
- To compare cafefilter's performance against existing popular filtering methods.
Main Methods:
- Large-scale study involving querying the PIR protein database.
- Application and analysis of popular query filtering techniques.
- Development and testing of the cafefilter technique, which considers motif distribution.
Main Results:
- Unselective application of popular filtering techniques can reduce retrieval effectiveness.
- Cafefilter demonstrates comparable or superior performance to existing tools in large-scale tests.
- The effectiveness of filtering is dependent on the query and database characteristics.
Conclusions:
- Selective filtering is necessary for optimal homology search performance.
- Cafefilter offers a robust alternative for masking low complexity regions.
- Database motif distribution is a key factor for effective query filtering.