用基于图形的注意力模型预测特定于背景的异形变异的变异解决预测
Aviya Litman1, Zhicheng Pan2, Ksenia Sokolova3
1Quantitative and Computational Biology Program, Princeton University, Princeton, NJ 08540, USA; Lewis-Sigler Institute for Integrative Genomics, Princeton University, Princeton, NJ 08540, USA.
Cell genomics
|January 17, 2026
概括
新的人工智能工具Otari分析全长基因转录,揭示遗传变化如何影响RNA拼接. 这有助于理解复杂的疾病,如自闭症,通过精确指明异型失调.
科学领域:
- 基因组学和分子生物学
- 计算生物学和生物信息学
背景情况:
- 细胞基因产生多个转录异型,对转录和蛋白质组多样性和功能调节至关重要.
- 遗传变异可以改变RNA处理信号,影响异构体结构和丰度,但在全长异构体分辨率下建模这些效应是复杂的.
研究的目的:
- 介绍Otari,一个基于注意力的图形神经网络框架,用于预测组织特异的差异化异形态丰度.
- 通过整合序列衍生信号来实现基因变异效应的异型解析解释.
主要方法:
- 奥塔里接受了人类基因组序列和长时间读取的转录组在各种人类组织和大脑区域的训练.
- 该框架整合了序列衍生的表观遗传和后转录信号,以预测异构体的丰富性.
- 奥塔里被应用于大型变体数据集,包括自闭症队列.
主要成果:
- 奥塔里成功地预测了组织特异差异化异型的丰富性.
- 该模型揭示了在基因水平上无法检测到的异形失调模式.
- 确定了与自闭症病理生理学相关的异形丰度和微埃克逊使用的变异驱动变化.
结论:
- 奥塔里提供了一个强大的资源,用于在多种组织中进行大规模的异构体水平分析.
- 该框架增强了对转录异型的遗传变异影响的解释.
- 在异形分辨率上,Otari有助于更深入地了解疾病机制,如自闭症.
相关概念视频
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
End Point Prediction: Gran Plot
1.2K
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
1.2K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Variability: Analysis
448
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
448
Conservation of Protein Domains Over Different Proteins
14.1K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.1K


