Related Experiment Video
Updated: Nov 29, 2025

Design and Optimization Strategies of a High-Performance Vented Box
Published on: June 9, 2023
High-Performance, Graphics Processing Unit-Accelerated Fock Build Algorithm
Giuseppe M J Barca1, Jorge L Galvez-Vallejo2, David L Poole2
1Research School of Computer Science, Australian National University, Canberra, Australian Capital Territory 2601, Australia.
Abstract:
We present a high-performance, GPU (graphics processing unit)-accelerated algorithm for building the Fock matrix. The algorithm is designed for efficient calculations on large molecular systems and uses a novel dynamic load balancing scheme that maximizes the GPU throughput and avoids thread divergence that could occur due to integral screening. Additionally, the code adopts a novel ERI digestion algorithm that exploits all forms of permutational symmetry, combines efficiently the evaluation of both Coulomb and exchange terms together, and eliminates explicit thread synchronization requirements. Performance results obtained using a number of large molecules reveal remarkable speedups up to 24.4× with respect to the QUICK GPU code and up to 237× with respect to the GAMESS CPU parallel code.
Related Concept Videos
Fast Fourier Transform
The computational efficiency of the FFT becomes...
Parallel Processing
Fast Decoupled and DC Powerflow
Acceleration Vectors
Accelerating Fluids
The motion of the liquid within this infinitesimal cylinder is considered to obtain the pressure difference. Three vertical forces act on this liquid:

