Related Experiment Video
Updated: Jan 24, 2026

08:35
An Operant Intra-/Extra-dimensional Set-shift Task for Mice
Published on: January 22, 2016
12.7K
Penalized logistic regression with low prevalence exposures beyond high dimensional settings
Sam Doerken1,2, Marta Avalos3,4, Emmanuel Lagarde3
1Institute of Medical Biometry and Statistics, Faculty of Medicine and Medical Center, University of Freiburg, Freiburg, Germany.
Plos One
|May 21, 2019
Summary
Estimating rare risk factors for binary outcomes is difficult. Penalized regression methods like Firth correction and boosting significantly improve estimation, especially for ultra-low prevalences below 0.1%.
Area of Science:
- Biostatistics
- Epidemiology
- Statistical modeling
Background:
- Estimating risk factors with very low exposure prevalence for binary outcomes presents challenges for standard statistical methods like logistic regression.
- Classical methods often yield unreliable results in these low-prevalence scenarios.
Purpose of the Study:
- To evaluate the effectiveness of penalized regression techniques in improving the estimation of low-prevalence risk factors.
- To compare Firth correction, ridge, lasso, and boosting methods in low-dimensional settings.
Main Methods:
- Utilized a large unmatched case-control study dataset (France, 2005-2008) on prescription medicines and road traffic accidents.
- Conducted an accompanying simulation study to assess method performance.
- Applied Firth correction, ridge, lasso, and boosting regression techniques.
Main Results:
- Penalized regression methods, particularly Firth correction and boosting, significantly enhance the estimation of risk factors with prevalences below 0.1%.
- These improvements are especially pronounced for ultra-low prevalences.
- The study demonstrated the utility of these methods even in low-dimensional settings.
Conclusions:
- Penalized regression techniques offer substantial benefits for estimating low-prevalence risk factors in binary outcome studies.
- Firth correction and boosting are recommended for improving risk factor estimation, especially at extremely low prevalences.
- The findings support the use of penalized techniques when dealing with a moderate number of low-prevalence exposures.
Related Concept Videos
Regression Toward the Mean
6.9K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.9K
Prevalence and Incidence
1.7K
In statistical epidemiology and health sciences, two essential metrics—prevalence and incidence—are fundamental for understanding disease dynamics within a population. These measures enable public health officials, epidemiologists, and researchers to assess the burden of diseases, allocate resources effectively, and design impactful public health policies and interventions.
Prevalence indicates the proportion of individuals in a population who have a specific disease or health...
Prevalence indicates the proportion of individuals in a population who have a specific disease or health...
1.7K
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Correlation and Regression
3.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.2K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K
Microsoft Excel: Regression Analysis
1.5K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
1.5K

