基于相似性的配对提高了罗神经网络在回归任务和不确定性量化方面的效率
Yumeng Zhang1,2, Janosch Menke3,4, Jiazhen He5
1Medicinal Chemistry, Research and Early Development, Respiratory and Immunology (R&I), BioPharmaceuticals R&D, AstraZeneca, 43183, Gothenburg, Sweden.
Journal of cheminformatics
|August 30, 2023
概括
本研究介绍了一种基于相似性的新配对方法,用于训练罗神经网络,提高回归任务的效率和预测准确性. 该方法减少了算法复杂性,并提高了对机器学习预测的信心.
科学领域:
- 计算化学是一种计算化学.
- 机器学习 机器学习
- 深度学习是一种深度学习.
背景情况:
- 罗网络是一种神经网络类别,有两个相同的子网络,共享权重,但接收不同的输入.
- 训练姆网络进行回归任务的传统方法通常涉及到详尽的配对,导致高计算复杂性.
研究的目的:
- 开发一种更有效的基于相似性的配对方法,用于训练罗神经网络进行回归任务.
- 在机器学习模型中提高预测性能和量化预测不确定性.
主要方法:
- 开发了一种新的基于相似性的配对方法,将算法复杂度从O(n^2) 降低到O(n).
- 一个带有圆形指纹的多层感知子被用作概念验证.
- 基于变压器的Chemformer被集成到语神经网络架构中.
- 用参考化合物的预测差异来测量预测不确定性.
主要成果:
- 基于相似性的配对方法在三个物理化学数据集中证明了更好的预测性能.
- 整合Chemformer使得从简化分子输入线路输入系统 (SMILES) 表示中能够提取特定任务的特征.
- 在高预测准确度和高信心 (低不确定性) 之间观察到强烈的相关性.
结论:
- 拟议的基于相似性的配对方法为训练语网络在回归任务中的效率和性能提供了显著的改善.
- 该方法提供了一种可靠的方法来评估预测不确定性,将准确性与信心联系起来.
- 这些发现突显了相似性质原理在推进机器学习应用中的实用性.
相关概念视频
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K
Wilcoxon Signed-Ranks Test for Matched Pairs
160
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
160
Survival Tree
109
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
109
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K
Per-Unit Sequence Models
93
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
93
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


