Related Experiment Video
Updated: Nov 28, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
855
Identification of Semantically Similar Sentences in Clinical Notes: Iterative Intermediate Training Using Multi-Task
Diwakar Mahajan1, Ananya Poddar1, Jennifer J Liang1
1IBM Research, Yorktown Heights, NY, United States.
JMIR Medical Informatics
|November 27, 2020
Summary
This study developed an iterative training approach for clinical text similarity, achieving state-of-the-art results in the 2019 ClinicalSTS challenge by effectively leveraging labeled data and advanced language models.
Area of Science:
- Natural Language Processing
- Clinical Informatics
- Machine Learning
Background:
- Electronic health records (EHRs) contain redundant information due to templated notes and copy-pasting.
- Effective utilization of EHR data is hindered by information redundancy.
- Measuring semantic similarity in clinical text is crucial for information condensation.
Purpose of the Study:
- To enhance semantic textual similarity performance in the clinical domain.
- To improve the robustness of clinical text similarity models.
- To leverage manually labeled data and contextualized embeddings for better results.
Main Methods:
- Utilized the ClinicalSTS dataset of 1642 annotated clinical text pairs.
- Developed an iterative intermediate training approach using multi-task learning (IIT-MTL).
- Applied IIT-MTL to ClinicalBERT and ensembled with BioBERT, MT-DNN, RoBERTa, and handcrafted features.
Main Results:
- Achieved state-of-the-art performance in the 2019 n2c2/OHNLP ClinicalSTS challenge.
- Ranked first among 87 submitted systems with a Pearson correlation coefficient of 0.9010.
- The winning system was an ensemble model incorporating IIT-MTL on ClinicalBERT, BioBERT, MT-DNN, and medication features.
Conclusions:
- IIT-MTL effectively utilizes annotated data from related tasks for improved performance on data-limited target tasks.
- This approach offers new possibilities for optimized dataset selection in the clinical domain.
- The study contributes to generating more robust and universal contextual text representations.
Related Concept Videos
Improving Translational Accuracy
12.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
12.9K
Improving Translational Accuracy
3.4K
3.4K
