Related Experiment Video
Updated: Aug 29, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Stigmatizing Language in Gender-Expansive Patient Records: Corpus Development, Disparity Analysis, and Natural
Liyang Xue1, Mary Chayko1, Vivek Kumar Singh1
1Department of Library and Information Science, Rutgers, The State University of New Jersey, New Brunswick, NJ, United States.
Background:
Stigmatizing language (SL) in electronic health records (EHRs) can influence clinical decision-making, propagate bias across care encounters, and undermine patient trust. Gender-expansive patients (GEPs) may be particularly vulnerable to documentation-based stigma; however, large-scale quantitative evidence and fairness-aware evaluation of automated SL detection methods remain limited.
Objective:
This study aims to construct a gender-expansive-inclusive EHR corpus, quantify demographic disparities in SL using parallel outcome definitions with and without misgendering, and evaluate fairness-aware natural language processing (NLP) methods for automated detection of SL.
Methods:
We developed an annotated corpus of 754 clinical notes from the MIMIC-IV (Medical Information Mart for Intensive Care IV) database, including 366 GEP notes and 388 matched nongender-expansive patient (NGEP) notes, labeled for SL and its subtypes. Parallel outcome definitions were constructed with and without misgendering. Multivariable logistic regression was used to assess associations between gender-expansive status and SL while adjusting for race, age, and primary language. Multiple NLP models were evaluated for SL detection. Fairness-aware post hoc threshold optimization based on equalized odds principles was applied using training-set predictions to reduce subgroup error disparities.
Results:
SL was identified in 62.3% (228/366) of GEP notes compared with 25.5% (99/388) of NGEP notes. When misgendering was excluded, the prevalence remained higher among gender-expansive notes at 41.8% (153/366). In multivariable models, gender-expansive status was strongly associated with stigmatizing documentation when misgendering was included (adjusted odds ratio [OR] 4.87, 95% CI 3.54-6.70) and remained significant when misgendering was excluded (adjusted OR 2.12, 95% CI 1.54-2.91). Post hoc equalized odds threshold optimization for a state-of-the-art transformer-based detector reduced the difference in false-positive rate (ΔFPR) from 15.76 to 6.65 percentage points (pp) and the difference in true-positive rate (ΔTPR) from 7.22 pp to substantially lower levels while maintaining similar accuracy (82.78%-83.44%). When misgendering was excluded, fairness optimization reduced ΔFPR to 2.98 pp and ΔTPR to 0.65 pp, with an overall accuracy of 88.08%.
Conclusions:
SL is common in EHR documentation and disproportionately affects GEPs, and automated detection models show persistent subgroup performance gaps. Disparities remained significant even when misgendering was excluded, indicating that bias extends beyond identity-specific errors to broader evaluative language. This study introduces the first annotated corpus focused on SL in GEP documentation, quantifies demographic disparities, and demonstrates practical fairness-aware NLP strategies that can reduce error-rate inequities while preserving accuracy. These findings support equity-focused interventions to address SL through fairness-aware models as assistive auditing tools with human oversight.