Related Experiment Video
Updated: Jun 28, 2025

09:27
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
10.0K
Variant Effect Prediction in the Age of Machine Learning
Yana Bromberg1,2, R Prabakaran3, Anowarul Kabir4
1Department of Biology, Emory University, Atlanta 30322, Georgia, USA yana.bromberg@emory.edu.
Cold Spring Harbor Perspectives in Biology
|April 15, 2024
Summary
Unsupervised deep learning methods show promise for analyzing genetic variants, matching or exceeding supervised approaches. These faster, unlabeled methods are ideal for large-scale evaluations, especially in nonhuman proteins.
Area of Science:
- Computational biology
- Genomics
- Bioinformatics
Background:
- Supervised methods for analyzing single amino acid substitutions are limited by small, curated datasets and inconsistent variant effect definitions.
- Deep learning (DL) presents an opportunity to analyze unannotated protein sequences, potentially overcoming limitations of supervised approaches.
Purpose of the Study:
- To evaluate the performance of unsupervised, deep learning-based methods against traditional supervised methods for predicting the impact of genetic variants.
- To assess the potential of machine learning to interpret the 'language of life' from protein sequences.
Main Methods:
- Comparative analysis of supervised and unsupervised (deep learning) computational methods.
- Evaluation of methods based on performance metrics and variant effect prediction types.
- Focus on single amino acid substitutions arising from single-nucleotide variants in coding regions.
Main Results:
- Some unsupervised methods perform comparably to or better than existing supervised methods.
- Unsupervised methods are computationally faster, enabling large-scale variant effect evaluations.
- Method performance varies significantly by evaluation metrics and the specific type of variant effect predicted.
Conclusions:
- Unsupervised deep learning methods offer a viable and efficient alternative for variant effect prediction.
- Further research and validation are needed, particularly for nonhuman proteins where unsupervised methods show significant promise.
Related Concept Videos
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Cause and Effect
10.9K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
10.9K
Variability: Analysis
141
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
141
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K

