分布式内存,GPU加速 Fock 构造用于混合,高斯基密度函数理论.
David B Williams-Young1, Andrey Asadchev2, Doru Thom Popovici1
1Applied Mathematics and Computational Research Division, Lawrence Berkeley National Laboratory, Berkeley, California 94720, USA.
本研究介绍了新的分布式内存算法,用于在图形处理单元 (GPU) 上进行电子结构计算. 这些方法提高了混合密度函数理论 (DFT) 对大型原子系统的计算的性能和可扩展性.
科学领域:
- 计算化学计算化学
- 材料科学 材料科学 材料科学
- 高性能计算 高性能计算
背景情况:
- 现代超级计算机越来越多地使用图形处理单元 (GPU) 来加速计算.
- 为大规模并行GPU架构优化电子结构方法是越来越多的研究重点.
- 对于高斯基原子轨道方法的现有GPU加速通常针对共享内存系统,限制了大规模并行性.
研究的目的:
- 开发和介绍分布式内存算法来评估库伦和精确交换矩阵在混合科恩-夏姆密度函数理论 (DFT).
- 为了在使用高斯基数集的大型系统上实现高效的电子结构计算.
- 为解决量子化学中大量并行GPU算法的需求.
主要方法:
- 实现分布式内存算法用于直接密度安装 (DF-J-Engine) 库伦矩阵评估.
- 开发分布式内存算法用于半数值 (sn-K) 精确交换矩阵评估.
- 使用了混合Kohn-Sham DFT与高斯基数组.
主要成果:
- 演示了开发的分布式内存算法的绝对性能.
- 展示了从数百原子到超过一千原子的系统的强大的可扩展性.
- 在Perlmutter超级计算机上最多在128个NVIDIA A100 GPU上验证了性能.
结论:
- 所介绍的分布式内存算法有效地利用了大量并行GPU资源进行电子结构计算.
- 这些方法在混合DFT中的大型原子系统中显示了显著的性能和可扩展性改进.
- 这项工作提升了现代超级计算机架构上大规模量子化学模拟的能力.
更多相关视频
08:04Excitonic Hamiltonians for Calculating Optical Absorption Spectra and Optoelectronic Properties of Molecular Aggregates and Solids
Published on: May 27, 2020
10:52Multiscale Sampling of a Heterogeneous Water/Metal Catalyst Interface using Density Functional Theory and Force-Field Molecular Dynamics
Published on: April 12, 2019
相关概念视频
Hybridization of Atomic Orbitals II
Hybridization of Atomic Orbitals I
Fermi Level Dynamics
Electron affinity in semiconductors refers to the energy gap between the minimum of its conduction band and the vacuum level and it is a critical parameter in determining how easily a semiconductor can accept additional electrons.
The work...
Gauss's Law
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Valence Bond Theory and Hybridized Orbitals
A σ bond (single bond in a Lewis structure) is a covalent bond in which the electron density is...
