Deploying and scaling distributed parallel deep neural networks on the Tianhe-3 prototype system.

Jia Wei1, Xingjun Zhang2, Zeyu Ji1

  • 1Xi'an Jiaotong University, Xi'an, 710049, Shaanxi, China.

Scientific Reports
|October 13, 2021
PubMed
Summary

Deep learning training on the Tianhe-3 supercomputer is accelerated using an optimized gradient synchronization strategy. This research enhances deep neural network (DNN) training performance on ARM-based architectures.