Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Generative Models and Sentence Transformers for the Recognition and Normalization of Continuous and Discontinuous
Areej Alhassan1,2, Viktor Schlegel1,3, Monira Aloud2
1Department of Computer Science, School of Engineering, University of Manchester, Manchester, United Kingdom.
This study introduces DiscHPO, a system for extracting genetic phenotypes from clinical notes. It effectively normalizes both continuous and discontinuous mentions, improving genetic condition representation in personalized healthcare.
Area of Science:
- Bioinformatics and Computational Biology
- Medical Informatics
- Genetics and Genomics
Background:
- Accurate extraction and normalization of genetic phenotypes from clinical reports are crucial for understanding genetic conditions and advancing personalized healthcare.
- Challenges exist in recognizing discontinuous entity spans within clinical text, hindering precise data interpretation.
- Dysmorphology relies heavily on accurate phenotype representation for diagnosis and research.
Purpose of the Study:
- To develop a system (DiscHPO) for accurate extraction and normalization of genetic phenotypes from clinical examination reports.
- To specifically address the challenge of identifying and normalizing discontinuous entity spans in dysmorphology assessments.
- To improve the representation of genetic conditions in personalized medicine through enhanced phenotype recognition.
Main Methods:
- A two-phase pipeline approach was employed, starting with a sequence-to-sequence model for named entity recognition (NER) of spans.
- An entity normalization phase utilized a sentence transformer biencoder for candidate generation and a cross-encoder reranker for concept selection.
- The system was evaluated within the context of the BioCreative VIII shared task, Track 3.
Main Results:
- The top-performing normalization model achieved an F1-score of 0.723, and the best span extraction model reached an F1-score of 0.665 on the test set.
- Both models outperformed baseline methods, demonstrating effectiveness in handling continuous and discontinuous spans.
- On the validation set, the system achieved an F1-score of 0.631 for exact matches on discontinuous spans.
Conclusions:
- Exact extraction of entity spans is not always required for successful phenotype normalization.
- Partial mention matches that capture essential concept information are sufficient for clinical utility.
- The DiscHPO system shows promise for downstream clinical applications requiring accurate genetic phenotype data.
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Improving Translational Accuracy
Improving Translational Accuracy
Background and Environment Affect Phenotype
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
Whole Body Regeneration
Automatic Processing and Automatic Social Behavior
