Related Experiment Video
Updated: Aug 14, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Understanding of Preeclampsia Risk Factors in Large Language Models Compared with a Validated Competing-Risks Model
Alexandra-Elena Cristofor1, Oriana-Maria Onicescu2, Denisa-Oana Zelinschi1
1Mother and Child Department, Grigore T. Popa University of Medicine and Pharmacy Iasi, 16 Universitatii Str, 700115 Iasi, Romania.
None:
Background: First-trimester screening for preterm preeclampsia relies on validated competing-risks models integrating maternal characteristics with biochemical and biophysical markers to generate individualized risk estimates. Large language models (LLMs) are increasingly used for pregnancy-related health information, yet their alignment with established clinical risk models remains unclear. Objective: to perform an exploratory, local perturbation-based assessment of how LLM-generated numerical risk estimates reproduce the direction and relative magnitude of established preeclampsia predictor effects compared with the Fetal Medicine Foundation (FMF) competing-risks model. Methods: A total of 129 synthetic clinical scenarios were generated from a low-risk reference pregnancy using a one-factor-at-a-time perturbation approach across 16 risk factors represented as 22 predictors. Eight LLMs (Claude, GPT, DeepSeek, Gemini, Copilot, Meta, Mistral, and Grok) estimated the probability of preeclampsia requiring delivery before 37 weeks. Corresponding risks were calculated using the FMF model. Model behavior was analyzed in logit space using a local perturbation-based modeling approach inspired by Local Interpretable Model-Agnostic Explanations (LIME) to derive feature-effect coefficients. Agreement with the FMF model was assessed using correlation, directional concordance, cosine similarity, and normalized root mean squared error, summarized using an exploratory composite score. Results: The FMF model identified mean arterial pressure, placental growth factor, parity, uterine artery pulsatility index, and chronic hypertension as dominant predictors. Alignment between LLM outputs and the FMF model was heterogeneous, with composite scores ranging from 0.59 to 0.82. Models with higher descriptive scores preserved predictor directionality (up to 90.9%) and rank ordering, but agreement in magnitude and scaling was limited (R2: 0.44-0.59). Intermediate models showed preserved directionality with reduced magnitude agreement, while lower-scoring models demonstrated more frequent sign inconsistencies and minimal variance explained. Conclusions: LLMs demonstrated partial, prompt-specific alignment with the FMF model in this local perturbation analysis, particularly for predictor direction and relative importance, but did not consistently reproduce quantitative effect sizes. This approach was intended to characterize local model behavior around a predefined reference case rather than evaluate clinically realistic combinations of interacting risk factors or global clinical prediction performance. Given the evolving nature of LLMs, ongoing reassessment using standardized approaches is required.
