Related Experiment Video
Updated: Jul 4, 2025

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
An evaluation of GPT models for phenotype concept recognition
Tudor Groza1,2,3,4, Harry Caufield5, Dylan Gration6
1Rare Care Centre, Perth Children's Hospital, 15 Hospital Avenue, Nedlands, WA, 6009, Australia. tudor.groza@health.wa.gov.au.
Large language models (LLMs) show promise for clinical phenotyping and phenotype annotation, with GPT-4 outperforming existing tools when using in-context learning. However, challenges remain due to LLM variability and cost.
Area of Science:
- Medical Informatics
- Computational Biology
- Natural Language Processing
Background:
- Clinical deep phenotyping and phenotype annotation are crucial for rare disease diagnosis and knowledge building.
- These processes traditionally use ontology concepts and machine learning for phenotype recognition.
- Large language models (LLMs) are increasingly used for Natural Language Processing (NLP) tasks.
Purpose of the Study:
- To evaluate the performance of the latest Generative Pre-trained Transformer (GPT) models, specifically ChatGPT (GPT-3.5-turbo and GPT-4.0), for clinical phenotyping and phenotype annotation.
- To compare LLM performance against established gold standard corpora and existing tools.
Main Methods:
- Utilized seven prompts of varying specificity with GPT-3.5-turbo and GPT-4.0 models.
- Evaluated performance on two gold standard corpora: publication abstracts and clinical observations.
- Employed in-context learning as a key experimental parameter.
Main Results:
- The best performance, using GPT-4.0 with in-context learning, achieved document-level F1 scores of 0.58 on abstracts and 0.75 on clinical observations.
- A mention-level F1 score of 0.7 surpassed current best-in-class tools.
- Performance significantly decreased without in-context learning.
Conclusions:
- GPT-4.0 demonstrates state-of-the-art performance for phenotype recognition when the task is constrained to known ontology subsets.
- Promising results are tempered by LLM non-determinism, high costs, and lack of run-to-run concordance, posing challenges for clinical application.
Related Concept Videos
Concepts and Prototypes
The brain organizes this information using concepts, which are mental categories grouping linguistic data,...
Background and Environment Affect Phenotype
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
Polygenic Traits
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Pedigree Analysis

