Predicting Bitcoin Prices Using Machine Learning
Athanasia Dimitriadou1, Andros Gregoriou2
1College of Business, Law and Social Sciences, University of Derby, Lonsdale House, Quaker Way, Derby DE1 3HD, UK.
Entropy (Basel, Switzerland)
|May 27, 2023
Summary
This study predicts Bitcoin price movements using machine learning. A logistic regression model achieved 66% accuracy, suggesting the Bitcoin market is not weak-form efficient.
Area of Science:
- Quantitative Finance
- Computational Economics
- Machine Learning Applications
Background:
- The efficient market hypothesis (EMH) suggests asset prices reflect all available information.
- Bitcoin's price volatility and unique characteristics present a challenge to traditional financial market efficiency.
- Predicting cryptocurrency movements requires robust analytical frameworks.
Purpose of the Study:
- To develop and evaluate machine learning models for Bitcoin price prediction.
- To assess the predictive power of various financial and macroeconomic variables on Bitcoin.
- To test the weak-form efficiency of the Bitcoin market.
Main Methods:
- Compiled a dataset of 24 explanatory variables from December 2014 to July 2019.
- Developed forecasting models using historical Bitcoin prices, cryptocurrencies, exchange rates, and macroeconomic data.
- Compared the performance of logistic regression, linear support vector machine, and random forest algorithms.
Main Results:
- The logistic regression model demonstrated superior performance, achieving 66% prediction accuracy.
- The model incorporating past Bitcoin values, other cryptocurrencies, exchange rates, and macroeconomic variables was effective.
- Outperformed linear support vector machine and random forest models in predicting Bitcoin movements.
Conclusions:
- The Bitcoin market exhibits characteristics that deviate from weak-form efficiency.
- Machine learning, particularly logistic regression, offers a viable approach for Bitcoin price forecasting.
- Further research can explore more complex models and additional variables for enhanced prediction accuracy.
More Related Videos
Related Concept Videos
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Microsoft Excel: Regression Analysis
720
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
720
Steps in Outbreak Investigation
155
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
155
Residuals and Least-Squares Property
7.5K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.5K
Expected Value
4.0K
The expected value is known as the "long-term" average or mean. This means that over the long term of experimenting over and over, you would expect this average. The expected average is represented by the symbol μ. It is calculated as follows:
4.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


