Related Experiment Videos

Analysis of patterns of violence against women in Brazil: an unsupervised machine learning approach

Andre Massahiro Shimaoka1, Antonio Carlos da Silva Junior1, José Marcio Duarte1

  • 1Departamento de Informática em Saúde, Escola Paulista de Medicina, Universidade Federal de São Paulo. R. Botucatu 740, Vila Clementino. 04023-062 São Paulo SP Brasil. andre.shimaoka@unifesp.br.

Analyze patterns of association between diagnoses related to domestic violence in hospital admissions and identify clinical-demographic profiles using unsupervised machine learning. Data from the SUS Hospital Information System between 2008 and 2023 were used, covering 90,798 hospitalizations of women aged 20 to 59 years with ICD-10 codes related to violence. Frameworks for data preparation were applied, as well as algorithms for identifying association rules between diagnoses and topic modeling. The hospitalization rate stabilized after 2012, with an average between 2.6 and 3.2 per 100,000 women. States such as São Paulo, Bahia, and Minas Gerais accounted for 46% of absolute cases; Rio Grande do Norte and Pará presented the highest proportional rates. The algorithm identified significant rules between types of injury and mechanisms of aggression. Latent Dirichlet Allocation modeling revealed nine distinct profiles, highlighting young women undergoing emergency surgeries due to polytrauma. The use of machine learning identified relevant clinical-epidemiological patterns to support prediction and surveillance strategies in Primary Care, contributing to addressing gender-based violence.

Related Concept Videos

Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
Statistical Analysis System (SAS)01:14

Statistical Analysis System (SAS)

SAS, short for Statistical Analysis System, is a powerful data analysis, management, and visualization tool. Developed by the SAS Institute in the early 1970s, SAS has evolved into a comprehensive software suite used across various industries for statistical analysis, business intelligence, and predictive modeling.
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Regression Analysis01:11

Regression Analysis

Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as: