利用长读组件和机器学习来增强短读可转换元件检测和基因型定型
Austin Daigle1,2, Logan S Whitehouse1,2, Roy Zhao3
1Department of Genetics, University of North Carolina, Chapel Hill, NC 27599.
bioRxiv : the preprint server for biology
|February 24, 2025
概括
可转移元素 (TE) 是基因组进化的关键. 我们的机器学习工具,TEforest,使用负担得起的短读序列,准确地检测TE插入,改进了现有的方法.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 分子进化分子进化
背景情况:
- 可转移元素 (TE) 是基因组进化至关重要的移动遗传序列.
- 长读数测序提高了TE检测的准确性,但成本高昂.
- 短读测序方法在从真实数据中准确检测TE方面存在局限性.
研究的目的:
- 开发一种机器学习方法 (TEforest),用于使用短读序列数据进行精确的TE插入和删除发现和基因型识别.
- 通过长读序列识别识别的TE作为机器学习模型的训练数据.
主要方法:
- 利用一个敏感的算法来识别从短读对齐中潜在的TE插入/删除站点.
- 从短读对齐中提取相关特征用于TE检测.
- 训练了一个随机森林模型,使用通过长读序列识别的TEs的基准真实数据集.
主要成果:
- 与传统方法相比,TEforest表现优越,识别了更多的真实阳性和更少的假阳性.
- 该方法准确地推断基因型和精确的插入断点在各种读取长度和覆盖范围.
- TEforest有效地学习TEs的短读签名,以前只有长读才能检测到.
结论:
- TEforest弥合了大规模人口研究和高精度长读组件之间的差距,用于TE分析.
- 这种用户友好的工具有助于研究TE流行率和全基因组的表型效应.
相关概念视频
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Non-LTR Retrotransposons
11.3K
As the name suggests, non-LTR retrotransposons lack the long terminal repeats characteristic of the LTR retrotransposons. Additionally, both LTR and non-LTR retrotransposons use distinct mechanisms of mobilization. Non-LTR retrotransposons are further divided into two classes - Long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs), both of which occur abundantly in most mammals, including humans. Some of the active non-LTR retrotransposons in humans are L1...
11.3K


