Related Experiment Video
Updated: May 3, 2026

19:57
An Affordable HIV-1 Drug Resistance Monitoring Method for Resource Limited Settings
Published on: March 30, 2014
19.2K
Sequencing artifacts in the type A influenza databases and attempts to correct them
David L Suarez1, Nikki Chester, Jason Hatfield
1Exotic and Emerging Avian Viral Disease Research Unit, Southeast Poultry Research Laboratory, Agricultural Research Service, USDA, Athens, GA, USA.
Influenza and Other Respiratory Viruses
|February 12, 2014
Summary
High school students identified and corrected errors in public influenza gene sequences. Over 22% of suspect sequences were fixed, highlighting the need for better data quality control in genetic databases.
Area of Science:
- Virology
- Bioinformatics
- Genetics
Background:
- Public databases contain over 276,000 influenza gene sequences.
- Sequence quality is dependent on the contributor.
Purpose of the Study:
- Identify and correct erroneous influenza gene sequences in public databases.
- Hypothesize that longer-than-expected gene sequences contain errors.
- Engage sequence submitters to rectify data inaccuracies.
Main Methods:
- Screened Type A influenza viruses for gene segments exceeding accepted lengths.
- Analyzed sequences with extra nucleotides at conserved non-coding ends.
- Investigated common error types including primer, plasmid, and polymerase-added sequences.
Main Results:
- Identified 1081 suspect influenza sequences.
- Common errors included residual primers, cloned plasmid DNA, and Taq polymerase artifacts.
- 22.8% (215) of suspect sequences were corrected in the first year.
- 138 new erroneous sequences were added in the second year.
Conclusions:
- Student project successfully identified and facilitated correction of numerous erroneous influenza sequences.
- Errors commonly stemmed from sequencing and cloning artifacts.
- Emphasized the critical need for enhanced data integrity in public genetic sequence databases.
Related Concept Videos
Viral Mutations
33.0K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
33.0K
Leaky Scanning
4.5K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
4.5K

