Related Experiment Videos
Odd odds interactions introduced through dichotomisation of continuous outcomes
Lutz P Breitling1, Hermann Brenner
1Division of Clinical Epidemiology and Aging Research, German Cancer Research Center, Bergheimer Str 20, D-69115 Heidelberg, Germany. l.breitling@dkfz-heidelberg.de
Journal of Epidemiology and Community Health
|August 21, 2009
Summary
Dichotomizing continuous variables can create misleading interactions between predictors, even with moderate sample sizes. This statistical artifact may compromise research validity, urging critical evaluation of data analysis methods.
Area of Science:
- Statistics
- Biostatistics
- Data Analysis
Background:
- Dichotomization of continuous variables is a common but criticized analytical practice.
- Its impact on outcome variables with multiple predictors remains an area of concern.
Purpose of the Study:
- To evaluate the effects of dichotomizing a continuous outcome variable when two predictors are involved.
- To specifically investigate the emergence and significance of interaction terms.
Main Methods:
- Simulated a log-normally distributed continuous outcome variable.
- Employed logistic regression analysis following dichotomization.
- Examined various cut-offs, predictor effects, dispersions, and interaction terms.
Main Results:
- Dichotomization introduced significant spurious interactions between predictor variables.
- These artificial interactions achieved statistical significance even with modest sample sizes.
- A real-world example demonstrated these issues using sex and weight predicting gamma-glutamyltransferase.
Conclusions:
- The study highlights a novel concern regarding the dichotomization of continuous variables.
- Researchers must critically assess potential dichotomization-induced artifacts impacting result validity.
Related Concept Videos
Odds Ratio
The odds ratio (OR) is a statistical measure used extensively in epidemiology and research to quantify the strength of association between exposure and outcome across different groups. Unlike relative risk, which compares the probabilities of an event occurring, the odds ratio compares the odds of an event occurring in the exposed group to the odds of it occurring in the unexposed group. The odds, in this context, are calculated as the probability of the event happening divided by the...
Introduction to Test of Independence
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
Contingency Table
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
Random Variables
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Bias
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Friedman Two-way Analysis of Variance by Ranks
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...