Related Experiment Video
Updated: Jul 22, 2026

08:27
Image Recognition and Parameter Analysis of Concrete Vibration State Based on Support Vector Machine
Published on: January 5, 2024
Predicting stock returns using machine learning combined with data envelopment analysis and automatic feature
Hoang Thanh Nhon1, Nga Do-Thi2, Thao Nguyen-Trang3,4
1Faculty of Commerce, Van Lang University, Ho Chi Minh City, Vietnam.
Plos One
|September 25, 2025
Summary
This study introduces business efficiency scores to predict stock returns, enhancing machine learning model accuracy. Combining these scores with automated feature engineering significantly improved predictions in the Vietnamese stock market.
Area of Science:
- Financial Markets
- Machine Learning
- Econometrics
Background:
- Predicting stock returns is crucial for investors.
- Existing methods often overlook business efficiency metrics.
- The Vietnamese stock market presents unique challenges for return prediction.
Purpose of the Study:
- To evaluate the efficacy of business efficiency scores in predicting stock returns.
- To compare the performance of various machine learning models for stock return prediction.
- To investigate the impact of automatic feature engineering on prediction accuracy.
Main Methods:
- Data Envelopment Analysis (DEA) to calculate business efficiency scores.
- Collection of financial data for 26 real estate firms (2019-2024).
- Application and comparison of machine learning models (e.g., Deep Neural Network, Gradient Boosted Tree) with technical, fundamental, and efficiency indicators.
Main Results:
- Business efficiency scores significantly improve stock return prediction accuracy.
- Deep Neural Network model showed reduced error metrics (RMSE, MAE, MAPE) with efficiency scores.
- Gradient Boosted Tree with efficiency scores and automated feature engineering achieved the best performance (MAE: 0.122, MAPE: 103.19).
Conclusions:
- Business efficiency scores are valuable predictors of stock returns.
- Automated feature engineering further enhances predictive power when combined with efficiency metrics.
- The findings offer a novel approach for improving investment strategies in emerging markets like Vietnam.
Related Concept Videos
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Econometric Views (EViews)
561
Econometric Views, often stylized as EViews, is a package that merges statistical analysis with econometric studies. It is designed to provide tools for time series analysis, forecasting, and econometric model simulation. The software originated from MicroTSP software and has evolved significantly since its inception in 1981. The history of EViews is marked by a continuous effort to enhance its computational speed and user interface. It was initially developed for large computing systems but...
561
Microsoft Excel: Regression Analysis
1.5K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
1.5K