相关实验视频
Updated: Jul 24, 2025

16:14
Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
13.6K
利用历史来预测分布式工作流中的罕见异常转移
Robin Shao1, Alex Sim2, Kesheng Wu2
1EECS, University of California at Berkeley, Berkeley, CA 94720, USA.
Sensors (Basel, Switzerland)
|July 8, 2023
概括
这项研究通过分析流量日志来预测科学计算中的缓慢网络连接. 低采样正常数据显著提高了机器学习模型的准确性,用于识别性能瓶.
科学领域:
- 科学计算是科学计算.
- 网络性能分析分析 网络性能分析
- 机器学习应用程序 机器学习应用程序
背景情况:
- 科学计算中的分布式数据密集型应用依赖于高效的数据共享.
- 缓慢的网络连接在分布式工作流中造成了重大瓶,影响了研究进展.
- 鉴定这些缓慢的连接是具有挑战性的,因为它们在维护良好的网络中频率低.
研究的目的:
- 开发一种用于识别科学计算环境中的缓慢网络连接的预测模型.
- 为了解决与检测罕见事件 (如数据传输缓慢) 固有的阶级不平衡问题.
- 评估分层采样技术在提高机器学习模型性能方面的有效性.
主要方法:
- 在国家能源研究科学计算中心 (NERSC) 分析了2021年1月至2022年8月的网络流量日志.
- 基于历史网络流量模式的特征工程,以识别低性能数据传输.
- 应用和比较各种分层采样技术,以减轻阶级不平衡.
- 使用均衡数据集的机器学习模型的培训和评估.
主要成果:
- 定义了一组基于历史记录的功能,用于检测缓慢的数据传输.
- 分层抽样,特别是正常连接的低抽样,在模型训练中被证明是非常有效的.
- 开发的模型在预测缓慢连接方面获得了0.926的高F1得分.
结论:
- 在科学计算中预测缓慢的网络连接是可行的,对于工作流的优化至关重要.
- 低采样正常数据是一种简单而强大的技术,可以克服这个领域的阶级不平衡.
- 这些发现提供了一种实际的方法来提高分布式科学计算的可靠性和效率.
相关概念视频
Steps in Outbreak Investigation
154
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
154
Distribution Reliability and Automation
134
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
134
Distributed Loads: Problem Solving
678
Beams are structural elements commonly employed in engineering applications requiring different load-carrying capacities. The first step in analyzing a beam under a distributed load is to simplify the problem by dividing the load into smaller regions, which allows one to consider each region separately and calculate the magnitude of the equivalent resultant load acting on each portion of the beam. The magnitude of the equivalent resultant load for each region can be determined by calculating...
678
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Unusual Results
3.2K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.2K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K

