Related Experiment Video
Updated: Aug 14, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Understanding of Preeclampsia Risk Factors in Large Language Models Compared with a Validated Competing-Risks Model
Alexandra-Elena Cristofor1, Oriana-Maria Onicescu2, Denisa-Oana Zelinschi1
1Mother and Child Department, Grigore T. Popa University of Medicine and Pharmacy Iasi, 16 Universitatii Str, 700115 Iasi, Romania.
Diagnostics (Basel, Switzerland)
|August 13, 2026
Summary
Large language models (LLMs) partially align with clinical models for predicting preterm preeclampsia risk, matching predictor direction but not always magnitude. Ongoing reassessment is needed as LLMs evolve.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Prediction Models
- Maternal Health
Background:
- First-trimester screening for preterm preeclampsia uses validated competing-risks models.
- Large language models (LLMs) are increasingly used for health information, but their clinical alignment is unclear.
- Established models integrate maternal characteristics, biochemical, and biophysical markers for risk estimates.
Purpose of the Study:
- To assess how LLM-generated risk estimates align with the Fetal Medicine Foundation (FMF) model.
- To evaluate LLM reproduction of predictor effects' direction and magnitude using local perturbation analysis.
- To compare eight different LLMs' performance against a gold-standard clinical model.
Main Methods:
- Generated 129 synthetic clinical scenarios from a low-risk reference pregnancy.
- Used a one-factor-at-a-time perturbation approach across 16 risk factors (22 predictors).
- Analyzed LLM (Claude, GPT, DeepSeek, Gemini, Copilot, Meta, Mistral, Grok) and FMF model outputs using local interpretable model-agnostic explanations (LIME)-inspired methods.
Main Results:
- FMF model identified mean arterial pressure, placental growth factor, parity, uterine artery pulsatility index, and chronic hypertension as dominant predictors.
- LLM-FMF model alignment was heterogeneous (composite scores 0.59-0.82), with better preservation of predictor directionality (up to 90.9%) than magnitude (R²: 0.44-0.59).
- Higher-scoring LLMs showed preserved direction and rank order but limited magnitude agreement; lower-scoring models had more sign inconsistencies.
Conclusions:
- LLMs demonstrated partial, prompt-specific alignment with the FMF model, particularly in predictor direction and relative importance.
- LLMs did not consistently reproduce quantitative effect sizes of preeclampsia risk factors.
- Ongoing reassessment of LLMs using standardized approaches is necessary due to their evolving nature.
