Related Experiment Videos
Neural network detects errors in the assignment of mRNA splice sites
S Brunak1, J Engelbrecht, S Knudsen
1Department of Structural Properties of Materials, Technical University of Denmark, Lyngby.
Nucleic Acids Research
|August 25, 1990
Summary
Neural networks identified errors in genetic databanks by analyzing pre-messenger RNA splicing signals. This highlights the need for computational checks to ensure data accuracy in genetic research databases.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Nucleotide sequence databanks are crucial for genetic research, but data reliability is a concern.
- Current error-detection methods for manually or electronically entered data in databanks like EMBL and GenBank are limited.
Purpose of the Study:
- To assess the reliability of human gene sequence data in public databanks.
- To develop computational methods for detecting errors in nucleotide sequence databanks.
Main Methods:
- Utilized neural networks trained on human gene sequences from EMBL and GenBank databanks.
- Focused on recognizing pre-messenger RNA (pre-mRNA) splicing signals within human genes.
- Investigated discrepancies found during the neural network training process.
Main Results:
- During training on 33 human genes from EMBL, seven genes disrupted the learning process, revealing errors.
- Further analysis of EMBL data identified discrepancies with original publications in three genes and wrongly assigned splicing frames in four genes.
- Training on 241 human sequences from GenBank uncovered nine new errors.
Conclusions:
- Errors in nucleotide sequence databanks can arise from typographical mistakes and misinterpretations of experimental data.
- Computer algorithms can effectively detect inconsistencies in genetic data before databank incorporation.
- Automated error detection is essential for maintaining the integrity of genetic research resources.