Related Experiment Video
Updated: Oct 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Incorporating Domain Knowledge Into Language Models by Using Graph Convolutional Networks for Assessing Semantic
David Chang1, Eric Lin2, Cynthia Brandt1,3,4
1Yale Center for Medical Informatics, Yale University, New Haven, CT, United States.
This study developed a natural language processing system to identify similar clinical text, reducing redundant information in electronic health records. The advanced model combines text and graph encoders, achieving high accuracy in assessing semantic similarity.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Machine Learning
Background:
- Electronic health record systems generate redundant clinical documentation via copy-paste functions.
- Identifying and removing similar text snippets is crucial for improving clinical summarization and documentation quality.
Purpose of the Study:
- To develop a natural language processing (NLP) system for assessing clinical semantic textual similarity.
- To assign similarity scores to pairs of clinical text snippets.
Main Methods:
- Leveraged BERT-based models as text encoders and graph convolutional networks (GCNs) as graph encoders.
- Incorporated linguistic and domain knowledge from the MedSTS dataset, including concept graphs.
- Explored data augmentation, ensembling, and knowledge distillation to enhance model performance.
Main Results:
- Achieved strong baseline performance using fine-tuned BERT and ClinicalBERT models (Pearson correlation coefficients up to 0.848).
- GCN-based graph encoders, especially with pretrained knowledge graph embeddings, boosted performance (r=0.868).
- Ensembling techniques, including multisource ensembling and knowledge distillation, further improved performance, reaching a Pearson correlation coefficient of 0.882.
Conclusions:
- A novel NLP system combining BERT and GCN encoders was developed for the MedSTS clinical semantic textual similarity task.
- The system effectively incorporates domain knowledge and demonstrates the potential of advanced language models in identifying redundant clinical information.
- Further research and dataset development are ongoing, but current results show promise for improving clinical documentation quality.
More Related Videos
Related Concept Videos
Modeling and Similitude
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Causes of Similarity-Dissimilarity Effect
Natural and Artificial Concepts

