Related Experiment Video
Updated: May 1, 2026

Characterization and Functional Prediction of Bacteria in Ovarian Tissues
Published on: October 23, 2021
Using large language models to identify prediagnostic clinical features of ovarian cancer from healthcare records: a
Garth Funston1, Namu Park2, Matthew Thompson3
1Centre for Cancer Screening, Prevention and Early Diagnosis, Wolfson Institute of Population Health, Barts and The London School of Medicine and Dentistry, Queen Mary University of London, London, UK g.funston@qmul.ac.uk.
Background:
Most women with ovarian cancer are diagnosed after developing symptoms. However, symptoms are often recorded as free text within electronic health records (EHRs), which is not readily accessible for research.
Aim:
To use EHRs to examine associations between coded and large language model (LLM)-extracted free-text clinical features with ovarian cancer diagnosis.
Design And Setting:
Population-based case-control study using EHRs and cancer registry data from women attending primary care, outpatient, and emergency clinics associated with the University of Washington, US.
Method:
In total, 136 women with ovarian cancer cases (diagnosed 2012-2019) were matched (age, clinic type) to 1360 control participants. Twelve months of prediagnosis coded and free-text data were extracted from EHRs. LLMs were tested on annotated notes, before extracting information on 17 prespecified clinical features. Univariate conditional logistic regression analyses were used to identify clinical features associated with ovarian cancer.
Results:
There were 14 clinical features that were more commonly identified from free text using LLMs than from codes in both the case and control groups. There were 14 features that were significantly associated with ovarian cancer when using codes and LLM-extracted data, but only eight features were significant using codes alone. Using both coded and LLM-extracted data, 11 features had odds ratios >2. Thirteen features were significantly associated when restricting analysis to early-stage (I-II) diagnosis.
Conclusion:
There is an identifiable ovarian cancer symptom signature within EHRs, with LLM-based natural language processing approaches enabling extraction of key non-coded symptom information. LLMs could support EHR-based research, while LLM-based clinical decision support tools may improve identification of patients with symptoms of possible cancer.
More Related Videos
09:08Integration of Bioinformatics Approaches and Experimental Validations to Understand the Role of Notch Signaling in Ovarian Cancer
Published on: January 12, 2020
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025