使用机器学习进行基因组预测:对合成和实证数据上的规范化回归,集合,基于实例和深度学习方法的性能进行比较
Vanda M Lourenço1, Joseph O Ogutu2, Rui A P Rodrigues3
1Center for Mathematics and Applications (NOVA Math) and Department of Mathematics, NOVA SST, 2829-516, Caparica, Portugal. vmml@fct.unl.pt.
BMC genomics
|February 7, 2024
概括
用于基因组预测的机器学习方法显示出不同的性能. 经典方法,如线性混合模型,由于效率和简单性,仍然是强大的竞争者.
科学领域:
- 基因组学就是基因组学.
- 机器学习 机器学习
- 生物信息学是一种生物信息学.
背景情况:
- 基因组选择使用分子标记器准确预测繁殖值.
- 机器学习方法越来越多地用于高维基因组数据.
- 对于基因组预测的机器学习组的比较研究很少见.
研究的目的:
- 对基因组预测的监督机器学习方法进行比较评估.
- 评估不同机器学习组的计算成本.
- 将机器学习方法与经典方法进行比较.
主要方法:
- 评估了规范回归,深度,集合和基于实例的学习算法.
- 使用一个模拟动物育种数据集和三个实证玉米数据集.
- 非正式评估预测性能和计算成本.
主要成果:
- 机器学习方法的性能和成本因数据集和特征而异.
- 规范化方法的复杂性越来越高,导致高计算成本,而无需保证准确度的提高.
- 经典的线性混合模型和规则化的回归方法显示出具有竞争力的性能,效率和简单性.
结论:
- 没有一种单一的机器学习方法在基因组预测方面普遍优越.
- 古典方法仍然是强大的竞争者,因为它们的性能和效率的平衡.
- 需要提高机器学习算法和资源的计算效率.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.5K
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
4.0K
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
End Point Prediction: Gran Plot
325
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
325
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Improving Translational Accuracy
10.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.5K
