在机器学习时代的变量效应预测
Yana Bromberg1,2, R Prabakaran3, Anowarul Kabir4
1Department of Biology, Emory University, Atlanta 30322, Georgia, USA yana.bromberg@emory.edu.
Cold Spring Harbor perspectives in biology
|April 15, 2024
概括
无监督深度学习方法显示出分析遗传变异的前景,匹配或超过监督方法. 这些更快,无标签的方法非常适合大规模的评估,特别是非人类蛋白质.
科学领域:
- 计算生物学是一种计算生物学.
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 对分析单氨基酸替代物的监督方法受到小,精心策划的数据集和不一致的变异效应定义的限制.
- 深度学习 (DL) 提供了分析未注释的蛋白质序列的机会,可能克服监督方法的局限性.
研究的目的:
- 评估基于深度学习的非监督方法的性能,与预测遗传变异影响的传统监督方法相比.
- 评估机器学习对从蛋白质序列中解释"生命语言"的潜力.
主要方法:
- 对监督和无监督 (深度学习) 计算方法的比较分析.
- 基于绩效指标和变异效应预测类型的方法的评估.
- 专注于编码区域中单核酸变异产生的单氨基酸替代.
主要成果:
- 一些无监督方法的性能与现有的监督方法相比或更好.
- 无监督方法在计算上更快,可以进行大规模的变异效应评估.
- 方法性能因评估指标和预测的特定类型变异效应而有很大差异.
结论:
- 无监督深度学习方法为变量效应预测提供了可行和高效的替代方案.
- 需要进一步的研究和验证,特别是对于非人类蛋白质,无监督的方法显示出显著的希望.
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Cause and Effect
10.9K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
10.9K
Variability: Analysis
141
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
141
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K


