Related Experiment Video
Updated: Jun 23, 2026

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
Published on: April 8, 2020
A compute unified system architecture for graphics clusters incorporating data locality.
Christoph Müller1, Steffen Frey, Magnus Strengert
1Visualisierungsinstitut der Universität Stuttgart, Stuttgart, Germany. christoph.mueller@vis.uni-stuttgart.de
This research introduces a development environment for distributed graphics processing unit (GPU) computing, enhancing scalability for multi-GPU systems and clusters. The system optimizes data locality and GPU utilization for faster execution in parallel processing.
Area of Science:
- Computer Science
- High-Performance Computing
- Parallel Computing
Background:
- Distributed GPU computing is crucial for accelerating complex computations.
- Existing systems often face scalability challenges in multi-GPU and cluster environments.
- Graphics Processing Units (GPUs) offer significant parallel processing capabilities.
Purpose of the Study:
- To present a development environment for distributed GPU computing.
- To extend the CUDA parallel programming model to higher levels of parallelism (PCI bus, network interconnects).
- To enhance scalability and performance in multi-GPU systems and graphics clusters.
Main Methods:
- Developed a system based on CUDA, extending its programming model.
- Implemented an extended API mimicking global memory across distribution layers.
- Introduced an automatic, data-locality-aware GPU-accelerated scheduling mechanism.
- Handled underlying communication mechanisms transparently for developers.
Main Results:
- The system effectively extends parallelism to PCI bus and network interconnects.
- The data-locality-aware scheduler significantly reduces transmitted data.
- Improved GPU utilization and faster execution times were observed.
- Demonstrated performance and scalability on multi-GPU systems and graphics clusters.
Conclusions:
- The presented development environment enhances distributed GPU computing.
- The system offers high scalability, particularly in network-interconnected environments.
- Automatic scheduling based on data locality is key to efficient distributed GPU computing.
Related Concept Videos
Parallel Processing
Multimachine Stability
In analyzing the system, the nodal equations represent the relationship between bus voltages, machine voltages, and machine currents. The nodal equation is given by:
Area Computation by the Alternative Coordinate Method
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
A Single-Component System
