Related Experiment Videos
Edge-cloud collaboration-driven predictive modeling for high-performance computing centers
Shuaiyin Ma1, Yeye Cao2, Yang Liu3
1School of Computer Science and Technology, Xi'an University of Posts and Telecommunications, Xi'an, 710121, China; Shaanxi Key Laboratory of Network Data Analysis and Intelligent Processing, Xi'an University of Posts and Telecommunications, Xi'an, 710121, China; Xi'an Key Laboratory of Big Data and Intelligent Computing, Xi'an University of Posts and Telecommunications, Xi'an, 710121, China; Shaanxi Union Research Centre of University and Enterprise for 5G+ Industrial Internet Communication, Xi'an University of Posts and Telecommunications, Xi'an, 710121, China.
Abstract:
Accurate coolant flow prediction is critical for active thermal management in high-performance computing (HPC) centers, yet it is inherently challenged by mixed-timescale dynamics and high-frequency workload surges. Existing deep learning methods often prioritize global accuracy on smoothed stationary trends, which may lead to phase delays during abrupt thermal transients. In addition, high-capacity architectures can introduce non-negligible computational overhead for latency-sensitive, resource-constrained edge controllers. To overcome these limitations, this study proposes a deployment-oriented edge-cloud collaboration (ECC) framework integrated with a transient-aware predictive architecture, named FS-Attention, designed to balance transient responsiveness, engineering deployability, and decision transparency. FS-Attention couples local feature synthesis, temporal-memory encoding, and attention-based temporal refinement to improve coolant-flow tracking under non-stationary operating conditions. Evaluations on the real-world Frontier supercomputer dataset show that the feature synthesis attention (FS-Attention) model achieves competitive full-year prediction accuracy, with a coefficient of determination (R²) of 0.8744 and a root mean square error (RMSE) of 0.0350. Under isolated critical thermal events (CTEs), FS-Attention obtains the lowest RMSE of 0.0757, slightly lower than the Temporal Fusion Transformer (TFT) and 5.61% lower than the standard Transformer. Platform-based profiling further shows an inference latency of 0.016 ms and a parameter size of 323.1 K, suggesting model-side compatibility with facility-side edge execution, while attention-shift analysis provides diagnostic evidence of temporally adaptive model behavior under dynamic thermal conditions.
Related Concept Videos
Parallel Processing
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Scale-Up Processes
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Distributed Loads
For example, consider a bookshelf filled with books stacked vertically adjacent to each other. The weight of the books is evenly distributed over the length of the shelf. As a result, the pressure at different locations on the surface of the...
Distribution Reliability and Automation