Determining the degree of randomness of descriptors in linear regression equations with respect to the data size

Michael C Hutter1

  • 1Center for Bioinformatics, Campus Building E2.1, Saarland University, 66123 Saarbrücken, Germany. michael.hutter@bioinformatik.uni-saarland.de

Related Concept Videos

Variation01:19

Variation

An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
Degrees of Freedom01:02

Degrees of Freedom

The degree of freedom for a particular statistical calculation is the number of values that are free to vary. Thus, the minimum number of independent numbers can specify a particular statistic. The degrees of freedom differ greatly depending on known and uncalculated statistical components.
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily assigned.
Degrees of Freedom01:02

Degrees of Freedom

The degree of freedom for a particular statistical calculation is the number of values that are free to vary. As a result, the minimum number of independent numbers can specify a particular statistic. The degrees of freedom differ greatly depending on known and uncalculated statistical components.
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily...
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Random Error01:04

Random Error

Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...