基于自我监督的对比性学习预测人类致病性开始损失变体
Jie Liu1, Henghui Fan1, Na Cheng2
1Information Materials and Intelligent Sensing Laboratory of Anhui Province, Institutes of Physical Science and Information Technology, Anhui University, Hefei, 230601, Anhui, China.
BMC biology
|August 9, 2025
概括
起始损失变异影响蛋白质生产,但很少有被分类. StartCLR使用自主监督学习来准确预测致病性启动损失变体,即使有有限的标记数据.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 分子遗传学 分子遗传学
背景情况:
- 起始损失变体破坏了翻译启动,影响了蛋白质的产生.
- 准确的病原性评估对于了解疾病机制和临床基因组学至关重要.
- 目前,由于数据限制,只有大约1%的人类启动损失变体被分类.
研究的目的:
- 开发一种新的计算方法,用于预测启动损失变体的病原性.
- 为了应对在分类遗传变异时有限的标记数据的挑战.
主要方法:
- 介绍了StartCLR,这是一个整合多种DNA语言模型嵌入的预测方法.
- 员工自主监督的预培训与监督的微调,以利用未标记和标记的数据.
- 利用对比学习来提高未标记数据的利用率.
主要成果:
- 在各种测试集中,StartCLR表现出强大的概括性和卓越的预测性能.
- 该方法有效地从多个维度捕捉了变体上下文信息.
- 即使仅在高可靠性标记数据上进行训练,StartCLR也保持或提高了预测准确性.
结论:
- 将自我监督的对比学习与未标记的数据相结合,有效地解决了标记的开始损失变体的稀缺问题.
- 开始CLR显示了改善病原性遗传变异分类的巨大潜力.
- 这种方法提高了基因组数据在临床实践中的实用性.
相关概念视频
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
End Point Prediction: Gran Plot
586
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
586
Single Nucleotide Polymorphisms-SNPs
15.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.9K
Survival Tree
159
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
159
Difference from Background: Limit of Detection
7.1K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
7.1K


