Related Experiment Videos
Predicting women's height from their socioeconomic status: A machine learning approach
Adel Daoud1, Rockli Kim1, S V Subramanian2
1Center for Population and Development Studies, Harvard T.H. Chan School of Public Health, Harvard University, United States.
Social Science & Medicine (1982)
|August 31, 2019
Summary
Socio-economic status (SES) explains little variance in women's height. Machine learning models showed negligible improvement over traditional regression, indicating limited predictive power and no significant non-linear relationships for social determinants of health research.
Area of Science:
- Public Health
- Socioeconomic Determinants of Health
- Human Welfare Indicators
Background:
- Social determinants of health literature frequently uses socio-economic status (SES) to explain variations in women's height, a key indicator of population welfare.
- Existing research often lacks a systematic evaluation of SES predictive power and potential non-linear relationships with women's height.
- A comprehensive understanding of SES's contribution to health outcomes is crucial for advancing social determinants of health research.
Purpose of the Study:
- To systematically evaluate the predictive power of socio-economic status (SES) on women's height.
- To investigate potential non-linear relationships between SES indicators (education, occupation, material wealth) and women's height.
- To compare the predictive performance of machine learning algorithms against traditional Ordinary Least Squares (OLS) regression.
Main Methods:
- Utilized Demographic and Health Surveys (DHS) data from 1,273,644 women across 66 low- and middle-income countries (1994-2016).
- Trained seven distinct machine learning algorithms and assessed their out-of-sample predictive power.
- Compared machine learning results with Ordinary Least Squares (OLS) regression, controlling for country, community, and sampling year fixed effects.
Main Results:
- In the OLS framework, SES explained only 0.7% of the total variance in women's height (R²).
- The best-performing machine learning model, a Bayesian neural network, offered a negligible 0.3% improvement in explained variance over OLS.
- Findings suggest no significant non-linear relationships between SES and women's height, highlighting the predictive limits of SES.
Conclusions:
- Socio-economic status (SES) has limited predictive power for women's height, with machine learning models showing minimal gains over traditional regression.
- The study indicates a lack of substantial non-linear associations between SES indicators and women's height.
- Scholars should report both the average effect and variance explained by SES to enhance understanding of social determinants of health.
Related Concept Videos
Polygenic Traits
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
Regression Toward the Mean
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Variation: Normal Distribution, Range, and Standard Deviation
In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data, such as ethnicity, can be tabulated into a frequency count to provide information about the proportion, as well as the variety of groups in a sample or population. On the other hand, researchers can perform a wider set of calculations on quantitative data. The mean, mode, and median, for instance, are central tendency measures to identify a...
Applications of Normal Distribution
The normal distribution is a useful statistical tool. One of its practical applications is determining the door height after considering the normal distribution of heights of persons, such that many can pass through it easily without striking their heads. The normal distribution can also determine the probability of a person having a height less than a specific height.
The heights of 15 to 18-year-old males from Chile from 1984 to 1985 followed a normal distribution. The mean height is 172.36...
The heights of 15 to 18-year-old males from Chile from 1984 to 1985 followed a normal distribution. The mean height is 172.36...