Related Experiment Videos
Evaluation of Glycine max mRNA clusters
1Biological Sciences Department, University of Missouri-Rolla, Rolla, MO, USA. rfrank@umr.edu
BMC Bioinformatics
|July 20, 2005
Summary
This study evaluated clustering algorithms for gene discovery in Glycine max (soybean). Results show that using multiple stringencies improves the accuracy of gene clusters, leading to a more reliable non-redundant gene set.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Clustering expressed sequence tags (ESTs) aids gene discovery and genome analysis.
- NCBI's UniGene is a widely used public resource for gene-oriented clusters.
- Assessing the accuracy of UniGene's clustering for species like soybean is crucial.
Purpose of the Study:
- To establish a non-redundant gene set for Glycine max by clustering mRNA sequences.
- To compare the accuracy of a novel algorithm with UniGene for gene clustering.
- To evaluate the impact of different stringency levels on clustering results.
Main Methods:
- Applied a clustering algorithm to Glycine max mRNA sequences using two stringency levels.
- Compared the algorithm's output with UniGene's clustering results.
- Analyzed discrepancies using nucleotide/amino acid alignments and author-reported gene classifications.
Main Results:
- Analysis of 12 clusters revealed instances of incorrect sequence grouping across all methods.
- Neither the novel algorithm (PECT) at two stringencies nor UniGene demonstrated significantly superior accuracy in separating paralogs.
- Discrepancies were observed in cluster composition among the three tested methods.
Conclusions:
- While all methods produced errors, employing multiple stringencies enhances the reliability of gene clusters.
- A sequential hierarchical approach with increasing stringencies can improve confidence in EST-only clusters.
- The findings suggest a need for refined clustering strategies to ensure accurate gene set representation.