Related Experiment Video
Updated: Jan 9, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Sociodemographic Bias in Large Language Model Clinical Trial Screening
Shelly Soffer1,2, Mahmud Omar3,4, Orly Efros2,5
1Institute of Hematology, Davidoff Cancer Center, Rabin Medical Center; Petah Tikva, Israel.
Background:
Large language models (LLMs) are increasingly used in randomized clinical trial (RCT) screening, but their potential for sociodemographic bias remains unclear.
Objective:
To determine whether LLM-based trial screening judgments vary with patient sociodemographic characteristics when clinical details and eligibility criteria are held constant.
Design Setting And Participants:
Cross-sectional evaluation of Phase II-III RCT protocols from ClinicalTrials.gov (U.S. adult populations; 2023-2024). For each protocol, we created 15 physician-validated clinical vignettes rendered in 34 versions: one control (no identifiers) and 33 identity variants spanning gender, race/ethnicity, socioeconomic status, homelessness, unemployment, and sexual orientation.
Exposures:
Identity labels applied to otherwise identical vignettes, evaluated by nine contemporary LLMs.
Main Outcomes And Measures:
Primary: eligibility domain score (1-5 Likert scale) comparing identity variants versus control. Secondary: adherence, resources, risk-benefit, and trust/attitude domains. Mixed-effects models estimated adjusted mean differences with multiplicity-corrected P values; differences <.10 considered trivial.
Results:
Of 69 protocols, 58 met inclusion criteria. Analysis of 5,324,400 model evaluations showed eligibility judgments were largely stable: most identity-related differences fell within ±0.05 (transgender woman -.008 [95% CI -.04 to .02]; White male .036 [.01 to .07]). Only homelessness exceeded the trivial threshold (-.121 [-.15 to -.09], P<.001). Secondary domains revealed socioeconomic gradients, particularly for adherence (homeless -.595, P<.001) and resources (homeless -.715, P<.001), with smaller trust/attitude effects and negligible risk-benefit differences.
Conclusions And Relevance:
Bias in LLM-assisted trial screening is conditional. Within fixed criteria, models reason consistently; outside them, they echo the inequities of their data. Responsible deployment in clinical research depends on preserving that boundary so that automation strengthens fairness in trial access rather than inheriting distortion.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Bias in Epidemiological Studies
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Surveys
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Confounding in Epidemiological Studies