Related Experiment Video
Updated: Mar 8, 2026

16:14
Trajectory Data Analyses for Pedestrian Space-time Activity Study
Published on: February 25, 2013
14.3K
Main Trend Extraction Based on Irregular Sampling Estimation and Its Application in Storage Volume of Internet Data
Beibei Miao1, Chao Dou2, Xuebo Jin3
1Baidu, Inc., Beijing 100085, China.
Computational Intelligence and Neuroscience
|January 17, 2017
Summary
This study presents a novel method for cleaning noisy internet data center storage volume time series. The approach accurately extracts trends, improving future storage volume predictions.
Area of Science:
- Computer Science
- Data Science
- Time Series Analysis
Background:
- Internet data center storage volume is a critical time series for business value.
- Real-world storage data is often "dirty," containing noise, missing values, and outliers, hindering accurate prediction.
- Extracting the main trend from such data is essential for reliable future forecasting.
Purpose of the Study:
- To develop and validate an irregular sampling estimation method for extracting the main trend of noisy time series data.
- To improve the accuracy of future storage volume predictions for internet data centers.
Main Methods:
- Utilized a Kalman filter to effectively remove noise, missing data, and outliers from the time series.
- Employed cubic spline interpolation and an averaging method to reconstruct the primary trend of the storage volume data.
- Applied the developed method to real-world internet data center storage volume data.
Main Results:
- The proposed method accurately estimates the main trend of internet data center storage volume time series.
- Experimental results demonstrate the effectiveness of the developed technique in handling "dirty" data.
- The method significantly contributes to more accurate future storage volume predictions.
Conclusions:
- The irregular sampling estimation method is a robust approach for trend extraction in noisy time series.
- Accurate trend extraction is crucial for reliable forecasting of internet data center storage volumes.
- This research offers a valuable contribution to time series analysis and data center management.
Related Concept Videos
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K
Sampling Methods: Overview
3.7K
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling.
In analytical chemistry, the choice of...
In analytical chemistry, the choice of...
3.7K
Estimation of the Physical Quantities
8.5K
On many occasions, physicists, other scientists, and engineers need to make estimates of a particular quantity. These are sometimes referred to as guesstimates, order-of-magnitude approximations, back-of-the-envelope calculations, or Fermi calculations. The physicist Enrico Fermi was famous for his ability to estimate various kinds of data with surprising precision. Estimating does not mean guessing a number or a formula at random. Instead, estimation means using prior experience and sound...
8.5K
Extraction: Partition and Distribution Coefficients
5.3K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
5.3K
Sampling Distribution
18.7K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
18.7K
Random Sampling Method
15.5K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
15.5K

