Related Experiment Video
Updated: Feb 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Bridging clinical narratives and structured phenotypes with large language models and sentence transformers
Jihao Cai1, Guozhuang Li1, Yongxin Yang2
1Department of Orthopaedic Surgery, State Key Laboratory of Complex Severe and Rare Diseases, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, Beijing 100730, China; Beijing Key of Big Data Innovation and Application for Skeletal Health Medical Care, Beijing, 100730, China; Key Laboratory of Big Data for Spinal Deformities, Chinese Academy of Medical Sciences, Beijing, 100730, China.
We developed LEAP (LLM-Enhanced Automated Phenotyping), a novel framework for extracting structured phenotypes from electronic health records. LEAP significantly improves the accuracy and reliability of phenotype data for genetic research and clinical applications.
Area of Science:
- Computational biology
- Medical informatics
- Genomics
Background:
- Structured phenotypes are crucial for Mendelian disorder diagnosis and genetic research.
- Electronic health records (EHRs) contain vast phenotypic data, but it is largely unstructured.
- Existing automated phenotyping methods struggle with semantic variability and contextual information in clinical narratives.
Purpose of the Study:
- To develop an advanced automated phenotyping framework addressing limitations of current deep learning models.
- To improve the extraction of standardized Human Phenotype Ontology (HPO) identifiers from unstructured clinical text.
- To enhance the utility of EHR data for genetic studies and clinical decision support.
Main Methods:
- Proposed LEAP, a two-stage framework combining a large language model (LLM) for phenotype extraction and a fine-tuned sentence-transformer for HPO mapping.
- The LLM component handles long clinical narratives without text chunking.
- The sentence-transformer model is trained on over 5.3 million instances for accurate and deterministic HPO identifier generation.
Main Results:
- LEAP demonstrated significant relative improvements in precision (19.68%-412.68%) and F1 score (44.14%-298.77%) on a real-world EHR test set compared to existing tools.
- Achieved robust performance on external benchmarks, validating its generalizability.
- The framework ensures the output of valid and deterministic HPO identifiers, overcoming LLM limitations.
Conclusions:
- LEAP offers a robust and accurate solution for automated phenotyping from unstructured EHR data.
- The framework enhances the standardization and usability of phenotypic data for downstream analyses, including gene prioritization.
- LEAP represents a significant advancement in leveraging AI for clinical informatics and precision medicine.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Stereotype Content Model
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
