应用到共享自行车系统的预测方法的可扩展性评估
Alexandra Cortez-Ordoñez1, Pere-Pau Vázquez2, José Antonio Sanchez-Espigares3
1Department of Statistics and Operations Research, UPC-BarcelonaTECH, Avda. Diagonal, 647, Planta 6, 08034 - Barcelona, Spain.
Heliyon
|October 9, 2023
概括
本研究评估了公共自行车共享系统 (BSS) 的预测算法. 先知和随机森林算法显示一致的结果,但小型BSS往往缺乏足够的数据来准确预测.
科学领域:
- 城市流动性 城市流动性
- 数据科学数据科学数据科学
- 运输工程 运输工程
背景情况:
- 公共自行车共享系统 (BSS) 在城市环境中越来越普遍.
- 预测分析对于BSS操作至关重要,包括需求预测和自行车再平衡.
- 当前的BSS算法评估往往缺乏跨不同系统大小的可扩展性评估.
研究的目的:
- 评估流行预测算法在不同尺寸的共享自行车系统 (BSS) 的性能.
- 为了确定算法,无论系统规模如何,提供一致的结果.
- 了解用于预测建模的较小BSS中的数据充足性挑战.
主要方法:
- 对成熟的预测算法进行评估.
- 在三个不同的BSS大小进行测试:小 (约20个站),中等 (400个站以上) 和大 (1500个站以上).
- 基于系统规模的算法准确性和可靠性的比较分析.
主要成果:
- 先知和随机森林在不同的BSS大小中展示了最一致的预测性能.
- 较小的BSS (约20个站) 经常显示不够的数据,阻碍了强大的算法性能.
- 算法的有效性受到可用的历史数据量的显著影响,这与系统大小相关.
结论:
- 由于其一致的性能,建议Prophet和Random Forest用于BSS预测任务.
- 数据可用性是BSS中成功预测建模的关键因素,特别是在较小的系统中.
- 未来的研究应该专注于改进数据稀缺的BSS环境中的预测方法.
相关概念视频
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Steps in Outbreak Investigation
139
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
139
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
53
Flood risk assessment involves careful planning and analysis to ensure the safety of communities near water retention structures. Capacity contours are a vital tool in this process, as they illustrate the potential spread of water at specific levels in a given area. In the context of building a bund across a small valley, these contours play a critical role in evaluating the safety of nearby residential areas.In this example, the bund is intended to store stormwater in the valley. The engineers...
53
Probability Histograms
11.7K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
11.7K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K


