Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Predicting Postpartum Hemorrhage Using Clinical Features Extracted With Large Language Models
Elizabeth G Woo1, Israel Zighelboim1, Tyler Gifford1
1Center for Computational Medicine and Clinical AI, Department of Medicine, and the Section of Ultrasound, Genetics, and the Fetal Neonatal Care Center, Department of Obstetrics and Gynecology, University of Chicago, Chicago, Illinois; the Department of Obstetrics and Gynecology and the Division of Gynecologic Oncology, St. Luke's Cancer Center, St. Luke's University Health Network, Bethlehem, Pennsylvania; the Department of Biomedical Data Science, Stanford University, Stanford, California; and Maternal Fetal Medicine, Oregon Health & Science University, Portland, Oregon.
Objective:
To evaluate whether large language models (LLMs) applied to prenatal clinical notes can predict postpartum hemorrhage (PPH) before the onset of labor and to compare model performance across outcome definitions, including a novel intervention-based definition.
Methods:
We conducted a retrospective cohort study within a large regional health network. Two outcome definitions for PPH were used: 1) estimated or quantitative blood loss (EBL-QBL) extracted from clinical notes; and 2) a clinical intervention-based PPH definition (cPPH) designed to capture significant hemorrhage requiring intervention, including transfusion, uterotonics, Bakri balloon, or hysterectomy. We evaluated three PPH prediction pipelines: 1) structured data only-supervised machine learning that used structured electronic medical record data; 2) LLM-direct-direct prediction that used a fine-tuned LLM applied to clinical notes; and 3) LLM-extract-interpretable models that used LLM-extracted features combined with structured data. Model performance was evaluated using an area under the receiver operating characteristic curve (AUROC) on a temporally held-out test set.
Results:
Among 19,992 deliveries, 1,156 patients (5.8%) met the EBL-QBL definition of PPH, 321 (1.6%) met the cPPH definition, and 309 (1.5%) met both definitions. The LLM-based direct prediction model achieved the highest AUROC for both PPH definitions (AUROC 0.79-0.80), followed by interpretable models that combined LLM-extracted features with structured data (AUROC 0.76-0.78). Models that used only structured data had the lowest AUROC (0.65-0.71). The LLM-extracted features approach identified 47 significant predictors, including established risk factors such as multiple gestation and previous cesarean delivery.
Conclusion:
These findings highlight the potential of LLM-based approaches to improve PPH risk stratification beyond structured data alone, with the feature extraction method offering a promising balance between predictive performance and clinical utility. Eventual integration of these methods into clinical workflows could improve early detection and guide targeted preventive interventions.
