Related Experiment Videos
The mutual information theory for the certification of rice coding sequences
Nicolas Carels1, Ramon Vidal, Ricardo Mansilla
1Laboratório de Bioinformática, Universidade Estadual de Santa Cruz, Rodovia Ilhéus/Itabuna km. 16, Ilhéus, Bahia, Brazil. carels@uesc.br
FEBS Letters
|June 16, 2004
Summary
Mutual information theory identified aberrant rice genes in databases, revealing about 10% of sequences had unusual features. This suggests potential biases in gene prediction algorithms used by researchers.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Accurate annotation of coding sequences is crucial for genomic research.
- Existing gene databases like GenBank and TIGR may contain sequences with aberrant features.
- Gene prediction programs are widely used but can be prone to biases.
Purpose of the Study:
- To apply mutual information theory for certifying annotated rice coding sequences.
- To identify and quantify genes with aberrant compositional features in public databases.
- To investigate potential biases in gene prediction algorithms.
Main Methods:
- Utilized mutual information theory for sequence analysis.
- Filtered coding sequences larger than 600 bp.
- Analyzed GC content trends (GC3% vs GC2%) for sequence classification.
- Compared identified aberrant sequences with published gene sets.
Main Results:
- Successfully screened out genes with aberrant compositional features from GenBank and TIGR datasets.
- Approximately 10% of rice coding sequences were identified as aberrant after redundancy cleaning.
- Rejected sequences exhibited distinct GC content trends compared to published gene sets.
Conclusions:
- Mutual information theory is effective for certifying genomic sequence quality.
- A significant proportion of annotated rice genes may possess aberrant features, potentially due to prediction algorithm biases.
- Findings highlight the need for improved gene prediction methods and data validation.