Related Experiment Video
Updated: May 30, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
The Social Construction of Categorical Data: Mixed Methods Approach to Assessing Data Features in Publicly Available
Theresa Willem1,2, Alessandro Wollek3, Theodor Cheslerean-Boghiu3
1Institute of History and Ethics in Medicine, School of Medicine and Health, Technical University of Munich, Munich, Germany.
Including categorical data in machine learning models can yield varied results. Analyzing data categories using mixed methods is crucial for equitable model development in healthcare.
Area of Science:
- Machine Learning
- Data Science
- Medical Informatics
Background:
- Computer scientists use categorical data (e.g., gender, skin color) to enhance machine learning models in data-scarce fields like healthcare.
- The impact of categorical data on model accuracy and equity in diverse populations remains underexplored.
Purpose of the Study:
- To investigate the influence of categorical data on machine learning model performance.
- To analyze data collection and publication processes affecting categorical data utility.
- To propose a mixed-methods approach for evaluating categorical data before machine learning training.
Main Methods:
- A mixed-methods approach was developed, grounded in the social construction of categories.
- Quantitative analysis assessed the impact of including/excluding categorical features on a transformer-based model using a Brazilian dermatological dataset (PAD-UFES 20).
- Qualitative analysis involved interviews with dataset authors to understand data collection and publication rationales.
Main Results:
- Quantitative analysis revealed inconsistent effects of categorical data on model predictions across different classes.
- Qualitative insights explained observed quantitative effects by highlighting the social construction and context-dependency of data categories.
- Findings underscore limitations of using publicly available datasets in contexts different from their origin.
Conclusions:
- Caution is advised when using categorical data from public datasets without considering their social construction and context.
- A social scientific, context-dependent analysis using mixed methods is recommended to assess categorical data utility for intended populations.
- This approach aids in ensuring equitable machine learning model development, especially in data-scarce domains.
More Related Videos
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Nominal Level of Measurement
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Bar Graph