Related Experiment Video
Updated: Jan 4, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.5K
Athena: Automated Tuning of k-mer based Genomic Error Correction Algorithms using Language Models
Mustafa Abdallah1, Ashraf Mahgoub1, Hany Ahmed2
1School of Electrical and Computer Engineering, Purdue University, West Lafayette, USA.
Scientific Reports
|November 8, 2019
Summary
Athena optimizes genome assembly by automatically tuning error-correction parameters using NLP language models. This approach improves sequencing accuracy and assembly quality without needing a reference genome.
Area of Science:
- Bioinformatics and Computational Biology
- Genomics and Next-Generation Sequencing (NGS)
- Natural Language Processing (NLP) applications
Background:
- Error-correction (EC) algorithms for genomics reads critically depend on optimal configuration parameters (e.g., k-mer value).
- Optimal parameters vary across datasets (species, platforms) and EC tools, necessitating adaptive tuning.
- Current methods often require a reference genome for ground truth, limiting applicability.
Purpose of the Study:
- To develop an adaptive method for automatically optimizing EC configuration parameters to enhance genome assembly quality.
- To leverage NLP language modeling techniques for efficient and accurate parameter tuning.
- To validate the use of NLP's 'perplexity' metric for quantifying EC performance.
Main Methods:
- Developed 'Athena', an algorithmic suite employing N-Gram and Recurrent Neural Network (RNN) language models from NLP.
- Repurposed the 'perplexity' metric from NLP to quantitatively assess EC performance.
- Utilized a hill-climbing search guided by perplexity to find optimal configuration parameters (e.g., k-value).
Main Results:
- Athena accurately identified optimal k-values for 7 real datasets across 3 EC algorithms (Lighter, Blue, Racer).
- Perplexity showed a strong negative correlation (≥73%) with EC quality across diverse datasets and error types.
- Athena's selected k-values achieved assembly alignment rates within 0.53% of brute-force optimization, improving NG50 by 4.72X.
Conclusions:
- NLP-based language modeling provides an effective and efficient approach for adaptive EC parameter optimization.
- The perplexity metric reliably predicts EC performance, enabling automated tuning without reference genomes.
- Athena significantly enhances genome assembly quality by optimizing EC parameters, demonstrating broad applicability.
Related Concept Videos
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Mismatch Repair
6.2K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
6.2K
Mismatch Repair
43.4K
Overview
43.4K
Proofreading
8.6K
Synthesis of new DNA molecules is carried out by the enzyme DNA polymerase, which adds nucleotides on the daughter strand complementary to the template DNA strand. DNA polymerase has a higher affinity to add the correct base and ensures fidelity during DNA replication. Furthermore, it exhibits proofreading activity during replication, using an exonuclease domain that cuts off incorrect nucleotides from the nascent DNA strand.
Errors During Replication are Corrected by the DNA Polymerase...
Errors During Replication are Corrected by the DNA Polymerase...
8.6K
Proofreading
59.6K
Overview
59.6K

