Related Experiment Video
Updated: Nov 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Predicting Semantic Similarity Between Clinical Sentence Pairs Using Transformer Models: Evaluation and
Mark Ormerod1, Jesús Martínez Del Rincón1, Barry Devereux1
1Institute of Electronics, Communications & Information Technology, School of Electronics, Electrical Engineering and Computer Science, Queen's University Belfast, Belfast, United Kingdom.
This study developed a clinical Natural Language Processing system for semantic textual similarity (STS) that achieved competitive results. Analysis revealed how transformer models represent semantic similarity, offering insights for future clinical NLP advancements.
Area of Science:
- Natural Language Processing (NLP)
- Clinical Informatics
- Machine Learning
Background:
- Semantic Textual Similarity (STS) is challenging in clinical text due to specialized language and abbreviations.
- Developing NLP systems for clinical STS requires robust models that understand nuanced medical terminology.
Purpose of the Study:
- To create an NLP system for predicting similarity scores between clinical sentence pairs.
- To analyze intermediate token vectors in transformer models to understand how semantic similarity is represented.
Main Methods:
- Averaged similarity scores from independently fine-tuned transformer models for clinical sentence pairs.
- Investigated model loss, decodability, and representational similarity of token vectors.
Main Results:
- Achieved a 0.87 correlation with ground-truth scores, ranking 6th out of 33 teams.
- Identified failure cases in modeling similarity for prescription details and overprediction with token overlap.
- Revealed divergent representational strategies between BERT and XLNet models.
Conclusions:
- Developed a competitive deep learning baseline for clinical STS without hand-crafted rules.
- Provided a detailed analysis of model outputs and learned biases in transformer models.
- Highlighted new research directions in model distillation and sentence embedding for clinical NLP.
Related Concept Videos
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
Improving Translational Accuracy
Improving Translational Accuracy
Transformers with Off-Nominal Turns Ratios
Energy Losses in Transformers
There are four main reasons for energy losses in transformers.
The first cause can be the high resistance of the...