Related Experiment Video
Updated: Mar 7, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
A data-centric approach to detecting and mitigating demographic bias in pediatric mental health text
Julia Ive1, Paulina Bondaronek2, Vishal Yadav3
1University College London, Institute of Health Informatics, London, UK. j.ive@ucl.ac.uk.
Background:
Healthcare Artificial Intelligence (AI) offers transformative potential but often inherits biases from training data, worsening disparities. While bias mitigation has focused on structured data, mental health relies on unstructured clinical notes, where linguistic differences and data sparsity pose challenges. This study aims to detect and reduce non-biological textual bias in AI models supporting pediatric mental health screening.
Methods:
We analyzed ~20,000 pediatric anxiety cases and matched controls (ages 5-15) from Cincinnati Children's Hospital records, where gender prevalence transitions from male-dominant in early childhood to female-dominant in adolescence. Anxiety prediction models were fine-tuned using a Transformer architecture optimized for computational efficiency. Classification parity across sex subgroups was evaluated, and we also verified that the model relied on clinically relevant words (using the LIME tool). Bias was mitigated through informative term filtering and systematic gender-biased text replacement.
Results:
Here, we show systematic under-diagnosis of female adolescents, with 4% lower accuracy and 9% higher false-negative rates compared to male patients. Notes for male patients are on average 500 words longer, and linguistic similarity metrics reveal distinct word distributions between sexes. Applying our de-biasing framework reduces diagnostic bias by up to 27%, improving equity in model performance.
Conclusions:
We develop and evaluate a data-centric de-biasing framework to address gender-based disparities in clinical text arising from non-biological differences, such as reporting practices and documentation styles. Our method selectively de-biases data by neutralizing biased language and normalizing information density while preserving clinically relevant content. Further validation across different models is essential before clinical deployment.
More Related Videos
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Regression Toward the Mean
Bias in Epidemiological Studies
Stereotypes, Prejudice, and Discrimination
Stereotype Content Model
Halo Effect
Diagnostic and Statistical Manual of Mental Disorders (DSM)