Related Experiment Video
Updated: May 26, 2025

07:28
Identification of Functionally-Relevant Lentivirus Integration Sites in an Insertional Mutagenesis Cell Library
Published on: January 10, 2025
202
Leveraging long-read assemblies and machine learning to enhance short-read transposable element detection and
Austin Daigle1,2, Logan S Whitehouse1,2, Roy Zhao3
1Department of Genetics, University of North Carolina, Chapel Hill, NC 27599.
Biorxiv : the Preprint Server for Biology
|February 24, 2025
Summary
Transposable elements (TEs) are key to genome evolution. Our machine learning tool, TEforest, accurately detects TE insertions using affordable short-read sequencing, improving upon existing methods.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Evolution
Background:
- Transposable elements (TEs) are mobile genetic sequences crucial for genome evolution.
- Long-read sequencing enhances TE detection accuracy but is costly.
- Short-read sequencing methods show limitations in accurately detecting TEs from real data.
Purpose of the Study:
- To develop a machine learning approach (TEforest) for accurate TE insertion and deletion discovery and genotyping using short-read sequencing data.
- To leverage TEs identified via long-read sequencing as training data for the machine learning model.
Main Methods:
- Utilized a sensitive algorithm to identify potential TE insertion/deletion sites from short-read alignments.
- Extracted relevant features from short-read alignments for TE detection.
- Trained a random forest model using a ground-truth dataset of TEs identified by long-read sequencing.
Main Results:
- TEforest demonstrated superior performance compared to traditional methods, identifying more true positives and fewer false positives.
- The method accurately infers genotypes and precise insertion breakpoints across various read lengths and coverages.
- TEforest effectively learns short-read signatures of TEs previously only detectable with long reads.
Conclusions:
- TEforest bridges the gap between large-scale population studies and high-accuracy long-read assemblies for TE analysis.
- This user-friendly tool facilitates the study of TE prevalence and phenotypic effects across genomes.
Related Concept Videos
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Non-LTR Retrotransposons
11.3K
As the name suggests, non-LTR retrotransposons lack the long terminal repeats characteristic of the LTR retrotransposons. Additionally, both LTR and non-LTR retrotransposons use distinct mechanisms of mobilization. Non-LTR retrotransposons are further divided into two classes - Long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs), both of which occur abundantly in most mammals, including humans. Some of the active non-LTR retrotransposons in humans are L1...
11.3K

