Related Experiment Video
Updated: Jul 25, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.1K
Assessing the Resilience of Machine Learning Classification Algorithms on SARS-CoV-2 Genome Sequences Generated with
Bikram Sahoo1, Sarwan Ali1, Pin-Yu Chen2
1Department of Computer Science, Georgia State University, Atlanta, GA 30303, USA.
Biomolecules
|June 28, 2023
Summary
Machine learning models effectively identify SARS-CoV-2 genomes despite sequencing errors. Spaced k-mers and weighted k-mers embedding methods improve accuracy for long-read sequencing data analysis.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Third-generation sequencing (TGS) generates long reads crucial for genome assembly, including SARS-CoV-2.
- High error rates in TGS data can compromise genome assembly and biological interpretation.
- Understanding SARS-CoV-2 evolution and transmission relies on accurate genome sequencing.
Purpose of the Study:
- To evaluate machine learning (ML) models for analyzing SARS-CoV-2 genome sequences with sequencing errors.
- To compare the effectiveness of six different embedding techniques for error-incorporated genome data.
- To identify robust ML approaches for accurate SARS-CoV-2 genome sequence identification.
Main Methods:
- Utilized ML models with six distinct embedding techniques.
- Analyzed SARS-CoV-2 genome sequences with simulated and random errors mimicking TGS profiles.
- Employed spaced k-mers and weighted k-mers embedding methods for vector generation.
Main Results:
- Spaced k-mers embedding demonstrated high accuracy in classifying error-free SARS-CoV-2 genomes.
- Spaced k-mers and weighted k-mers embedding methods showed high accuracy in predicting error-incorporated sequences.
- Fixed-length vectors generated by these methods contributed to high model performance.
Conclusions:
- The study highlights the efficacy of specific ML embedding techniques for handling TGS errors in SARS-CoV-2 genomes.
- Researchers can leverage these findings to select appropriate ML models for accurate genomic analysis.
- This work provides insights into improving the identification of critical SARS-CoV-2 genome sequences.
Keywords:
classificationembedding methodslong readmachine learningsequencing errorthird-generation single-molecule sequencing (TGS)More Related Videos
Related Concept Videos
Survival Tree
117
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
117
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Viral Mutations
32.5K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.5K
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
Leaky Scanning
5.2K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Gene Evolution - Fast or Slow?
7.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.2K

