Related Experiment Video
Updated: Jun 6, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Extraction of Human Phenotype Ontology (HPO) Concepts from Clinical Notes Utilizing Large Language Models (LLM) with
Michael Larsen1, Ian M Campbell2, Lori A Orlando3
1Division of Medical Genetics, University of Utah Health, Salt Lake City, UT, USA.
Real-time ontology grounding significantly enhances Human Phenotype Ontology (HPO) term extraction from clinical notes using large language models (LLMs). This approach improves accuracy and reduces errors, advancing genetic diagnosis and variant prioritization.
Area of Science:
- Genomics
- Artificial Intelligence
- Clinical Informatics
Background:
- Accurate Human Phenotype Ontology (HPO) term extraction from clinical notes is crucial for genetic diagnosis and variant prioritization.
- Large Language Models (LLMs) face challenges in balancing precision, hallucination avoidance, and ontology mapping accuracy.
- Retrieval-based grounding has shown promise in improving individual LLM performance.
Purpose of the Study:
- To evaluate the impact of real-time ontology grounding using external tools on HPO extraction metrics across diverse LLMs.
- To assess the Model Context Protocol (MCP) as a standardized, vendor-agnostic framework for integrating these tools.
- To compare the performance of tool-augmented LLMs against a commercial Electronic Health Record (EHR)-based HPO extraction tool.
Main Methods:
- Five LLMs (Claude Sonnet 4.5, GPT-5.1, Gemini 2.5 Pro, Grok 4.1, Qwen3 30B) extracted HPO terms from synthetic clinical genetics notes.
- Two conditions were tested: baseline (internal knowledge only) and tool-augmented (real-time HPO retrieval via MCP or native interfaces).
- Performance was measured by Precision, Recall, and F1-score, with manual adjudication for errors and hallucinations.
Main Results:
- Tool augmentation significantly improved mean aggregate F1-score from 0.46 to 0.72 (p < 0.001).
- Mapping Error Rate decreased from 40.9% to 7.8% (p < 0.001), and Precision increased from 56% to 90%.
- Tool-augmented LLMs outperformed the commercial benchmark in F1-scores and recall for inferred phenotypes.
Conclusions:
- Real-time ontology grounding substantially enhances HPO extraction across LLMs by minimizing mapping errors and improving phenotype inference.
- The Model Context Protocol offers a standardized, interoperable solution for deploying clinical LLM pipelines in genomic medicine.
- This approach supports reproducible and vendor-agnostic AI applications in healthcare.
Related Concept Videos
Methods of Documentation II: POMR
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll
