通过使用深度自动编码器检测异常来预测SARS-CoV-2血统的主导地位
Simone Rancati1, Giovanna Nicora1, Mattia Prosperi2,3
1Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Pavia, Italy.
bioRxiv : the preprint server for biology
|November 14, 2023
概括
DeepAutoCoV是一种新型的深度学习系统,可以预测未来的主导性严重急性呼吸综合征冠状病毒2 (SARS-CoV-2) 血统的高精度和早期预警时间. 这一进步有助于预防性公共卫生策略对抗不断演变的COVID-19变种.
科学领域:
- 病毒学 病毒学
- 基因组学就是基因组学.
- 计算生物学 计算生物学
背景情况:
- COVID-19 流行病的标志是 SARS-CoV-2 变种的持续出现,具有增强的传染性和免疫逃避.
- 预测新的主导血统的兴起对于有效的公共卫生反应至关重要.
研究的目的:
- 开发和验证DeepAutoCoV,这是一个无监督的深度学习系统,用于预测SARS-CoV-2的未来主导系 (FDL).
- 与基线方法相比,评估系统在早期检测和预测准确度方面的表现.
主要方法:
- 在超过1600万个SARS-CoV-2尖端蛋白序列上训练了一种无监督的深度学习异常检测系统 (DeepAutoCoV).
- 定义FDL为 (子) 血统,占每周GISAID提交的>10%.
- 通过使用全球和国家特定数据集验证了该系统,数据集覆盖了大约4年.
主要成果:
- DeepAutoCoV成功地在非常低的频率 (0.01%-3%) 上识别了FDL,中位数的预测时间为4-17周.
- 该系统表现出卓越的预测能力,比基线方法好大约5倍和25倍.
- 确定了与增加病毒适应性相关的特定突变,提供了可解释的见解.
结论:
- DeepAutoCoV提供了一个强大的工具,用于早期检测和预测新出现的SARS-CoV-2血统.
- 该系统能够提前显著地标记像B.1.617.2这样的变种,从而支持积极的公共卫生干预.
- 来自DeepAutoCoV的可解释输出可以指导针对不断发展的病毒威胁的预防策略的优化.
相关概念视频
Steps in Outbreak Investigation
135
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
135
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Classification of Signals
482
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
482
End Point Prediction: Gran Plot
344
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
344
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


