深度学习分子嵌入的数据融合用于属性预测
Robert J Appleton1, Brian C Barnes2, Alejandro Strachan1
1School of Materials Engineering and Birck Nanotechnology Center, Purdue University, West Lafayette, Indiana 47907, United States.
Journal of chemical information and modeling
|October 27, 2025
概括
这项研究引入了一种新的深度学习方法,用于使用来自单任务模型的融合嵌入器进行材料性质预测. 这种方法提高了准确性和效率,特别是在稀疏的数据中,优于标准的多任务学习模型.
科学领域:
- 计算化学是一种计算化学.
- 材料科学是一种材料科学.
- 机器学习 机器学习
背景情况:
- 深度学习模型提供准确的物质属性预测,但在稀疏的数据上扎.
- 现有的多任务学习方法受到数据完整性和属性相关性的限制.
- 稀少的数据集阻碍了材料科学中可靠的预测模型的开发.
研究的目的:
- 开发一个改进的多任务学习框架,用于使用稀疏数据进行物质属性预测.
- 在训练数据有限的情况下,提高预测模型的准确性和效率.
- 创建一种方法,利用预训练的单任务模型的属性特定表示.
主要方法:
- 从独立的,预训练的单任务模型中融合深度学习嵌入式.
- 开发一种新的多任务模型架构,重复使用这些嵌入式.
- 验证量子化学基准数据集和稀疏的实验数据的方法.
主要成果:
- 合并多任务模型在稀疏的数据集上显著优于标准多任务模型.
- 重复使用预训练的嵌入导致更丰富,属性特定的表示.
- 拟议的方法需要较少的可训练参数来进行模型扩展.
结论:
- 融合式嵌入提供了一个强大的解决方案,用于多任务学习与稀疏的数据在材料科学.
- 这种技术提高了预测准确性和模型效率,扩大了适用性.
- 该方法为更有效的数据驱动材料发现提供了基础.
相关概念视频
Predicting Molecular Geometry
44.7K
VSEPR Theory for Determination of Electron Pair Geometries
44.7K
Tagging and Fusion Proteins
8.3K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.3K
Molecular Models
43.5K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
43.5K
Prediction Intervals
3.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.2K
Molecular Geometry and Dipole Moments
18.0K
The VSEPR theory can be used to determine the electron pair geometries and molecular structures as follows:
18.0K
DNA Microarrays
20.6K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
20.6K


