Related Experiment Video
Updated: Oct 9, 2026

Using an Automated Hirschberg Test App to Evaluate Ocular Alignment
Published on: March 24, 2020
Benchmarking Generative Artificial Intelligence for the Analysis of Racial and Gender Disparities in Academic
Mark Ashamalla1, Stephanie Quon1, Faisal Khosa2
1Faculty of Medicine, University of British Columbia, Vancouver, British Columbia, Canada.
Objectives:
Generative artificial intelligence (AI) is increasingly being used to analyze and interpret health disparities data, but its accuracy, analytical depth, and ability to recognize intersectional and structural inequities remain uncertain. This benchmarking study evaluated the performance of ChatGPT (GPT-5.2), Google Gemini 3, and Claude Sonnet 4.5 in analyzing an existing dataset on racial and gender representation in US academic ophthalmology.
Methods:
Publicly available Association of American Medical Colleges (AAMC) Faculty Roster data for ophthalmology from 1966 to 2021 were analysed. No human subjects were recruited or randomized. Each AI model was prompted using a standardized query template to analyse demographic trends in academic rank, gender distribution, racial and ethnic composition, tenure status, and leadership representation. Outputs were systematically evaluated using a structured rubric assessing accuracy, interpretive depth, recognition of bias and intersectionality, statistical rigor, and recommendations, using the benchmark human-led study by Tao et al. (2024).
Results:
All three AI models identified the major trends previously reported in the benchmark study, including substantial faculty growth, increasing women representation, Asian faculty expansion, and persistent underrepresentation of groups underrepresented in medicine. ChatGPT provided the greatest statistical specificity, identifying women representation increasing by + 0.65 percentage points per year and women department chairs increasing by + 0.29 percentage points annually. Gemini emphasized structural tenure shifts, while Claude provided strong qualitative interpretation. Although Claude performed a race-by-gender analysis and ChatGPT examined gender representation by academic rank, none reproduced the specific intersectional framework and subgroup comparisons employed in the benchmark study.
Conclusions:
When benchmarked against a published human-led analysis, the three generative AI platforms reproduced broad demographic trends but varied in statistical rigour and interpretive depth. However, critical human oversight remains essential for nuanced intersectional and policy-relevant interpretation.
