Unsupervised Learning Applied to the Stratification of Preterm Birth Risk in Brazil with Socioeconomic Data

Márcio L B Lopes1, Raquel de M Barbosa2, Marcelo A C Fernandes3

  • 1Laboratory of Machine Learning and Intelligent Instrumentation, Federal University of Rio Grande do Norte, Natal 59078-970, Brazil.

Insights

Unsupervised learning identified socioeconomic factors linked to preterm birth (PTB) risk in Brazil. Municipalities with lower education and public services showed higher PTB rates, particularly in the North and Northeast regions.

Area of Science:

  • Public Health
  • Data Science
  • Socioeconomics

Background:

  • Preterm birth (PTB) poses significant risks to newborns, with multifactorial causes not fully understood.
  • Socioeconomic factors are recognized contributors to PTB risk, necessitating further investigation.

Purpose of the Study:

  • To stratify preterm birth risk in Brazil using unsupervised learning techniques.
  • To analyze the association between socioeconomic indicators and PTB occurrence at the municipal level.

Main Methods:

  • Generation of a novel dataset combining municipal socioeconomic data and PTB rates from Brazilian Federal Government sources.
  • Application of unsupervised learning algorithms including k-means, Principal Component Analysis (PCA), and DBSCAN for risk stratification.
  • Validation of identified clusters.

Main Results:

  • Discovery of four distinct clusters with high PTB occurrence and three with low PTB occurrence.
  • High PTB clusters characterized by lower educational attainment, poorer public services (sanitation, waste management), and a lower proportion of white population.
  • Geographic concentration of high PTB clusters in Brazil's North and Northeast regions.

Conclusions:

  • Socioeconomic status and the quality of public services significantly influence preterm birth risk.
  • Targeted interventions addressing education and public services may reduce PTB rates in vulnerable Brazilian municipalities.

Related Concept Videos

Stratified Sampling Method01:16

Stratified Sampling Method

Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
13.0K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
569
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
727
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.3K
Longitudinal Studies01:26

Longitudinal Studies

Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
262