Related Experiment Video
Updated: Jun 5, 2025

A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
Multiple imputation methods: a case study of daily gold price.
Ala Alrawajfi1,2, Mohd Tahir Ismail1, Sadam Al Wadi3
1School of Mathematical Science, Universiti Sains Malaysia, Penang, Penang, Malaysia.
This study evaluated data imputation methods for missing financial time-series data, finding k-nearest neighbor (KNN) imputation most accurate for gold prices. Performance declined with increased missing data proportions.
Area of Science:
- Financial econometrics
- Data science
- Statistical modeling
Background:
- Missing values are a common challenge in financial time-series data collection.
- Accurate data imputation is crucial for reliable financial analysis and forecasting.
Purpose of the Study:
- To compare the effectiveness of various data imputation techniques for financial time-series data.
- To identify the most robust imputation method for daily gold price data.
- To assess the impact of missing data proportion on imputation accuracy.
Main Methods:
- Evaluation of mean imputation, k-nearest neighbor (KNN), hot deck, random forest, support vector machine (SVM), and spline imputation.
- Validation using actual daily gold closing prices.
- Performance assessment using metrics: mean error (ME), mean absolute error (MAE), root mean square error (RMSE), mean percentage error (MPE), and mean absolute percentage error (MAPE).
Main Results:
- K-nearest neighbor (KNN) imputation demonstrated superior performance across all evaluated accuracy metrics.
- The predictive accuracy of all tested imputation methods decreased as the percentage of missing data increased.
- Spline imputation and random forest showed moderate performance, while mean imputation and hot deck were less effective.
Conclusions:
- K-nearest neighbor (KNN) is recommended as a highly effective and dependable method for imputing missing values in financial time-series data.
- The proportion of missing data significantly impacts the precision of imputation techniques.
- Further research could explore hybrid imputation methods or advanced machine learning models for complex financial datasets.
Related Concept Videos
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Regression Toward the Mean
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Censoring Survival Data
Precipitation Gravimetry
In determining nickel by gravimetric analysis, a precipitant of ethanolic dimethylglyoxime is added to a hot nickel salt solution. This is quickly followed by the dropwise addition of dilute ammonia solution until precipitation occurs. A...

