Prediction methods for babies' birth weight using linear and nonlinear regression analysis

Ilker Etikan1, Musa Kazim Caglar

  • 1Department of Biostatistics, Faculty of Medicine, Gaziosmanpasa University, Kislayolu üzeri, 60100 Tokat, Turkey. ietikan@gop.edu.tr

The aim of this study is to determine more accurate prediction methods between linear and non-linear methods for prediction of babies' birth weight among maternal demographic characteristics. Three hundred pregnant women were included in the study. Blood glucose level before and after ingestion of glucose load, age, body mass index, % of change in weight during pregnancy, height, gestational age, parity, and fetal sex were collected as independent variables and baby birth weight as dependent variable. In linear regression, least squares estimation method was used to estimate parameters. Non-linear regression method was performed using neural network model with multilayer perceptrons, back propagation method was preferred as learning algorithm. Coefficient of determination, R2, of the linear regression equation was found 59.8% and the standard error of the estimate was calculated as 325.69 gr. In non-linear regression method R2 value was also found 59.8% and standard error of estimate was calculated as 320.30 gr. According to the results of the present study, one method is not significantly better than the other. When "accuracy in prediction" is aimed, it is better to use the two methods and compare the results, and then decide on the selection of the favourable method.

Related Concept Videos

Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Microsoft Excel: Regression Analysis01:18

Microsoft Excel: Regression Analysis

Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
Regression Analysis01:11

Regression Analysis

Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Correlation and Regression00:53

Correlation and Regression

In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a negative...