Related Experiment Video
Updated: Oct 8, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking Effectiveness and Efficiency of Deep Learning Models for Semantic Textual Similarity in the Clinical
Qingyu Chen1, Alex Rankine1,2, Yifan Peng1,3
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD, United States.
Deep learning models show high effectiveness for semantic textual similarity (STS) but vary in efficiency and robustness. BERT models struggle with nuanced sentence pairs and are significantly slower, posing challenges for real-time clinical applications.
Area of Science:
- Natural Language Processing (NLP)
- Computational Linguistics
- Biomedical Informatics
Background:
- Semantic Textual Similarity (STS) measures sentence relatedness, crucial for clinical text analysis.
- The Open Health Natural Language Processing (OHNLP) Consortium developed an STS dataset, prompting research in clinical NLP.
- Deep learning (DL) models show promise but require thorough evaluation for real-world deployment.
Purpose of the Study:
- To benchmark the effectiveness and efficiency of top-ranked DL models for STS in a clinical context.
- To quantify the robustness and inference times of these models for real-time application validation.
- To address concerns regarding the practical utility of highly correlated DL models in production systems.
Main Methods:
- Benchmarked five DL models: CNN, BioSentVec, BioBERT, BlueBERT, and ClinicalBERT, plus a random forest baseline.
- Conducted 10 repetitions using official training and testing sets, reporting 95% CIs for Pearson correlation and running time.
- Evaluated additional metrics including Spearman correlation, R², and mean squared error for comprehensive analysis.
Main Results:
- All models demonstrated high effectiveness, with BioSentVec and BioBERT achieving the highest Pearson correlations (0.8497, 0.8481).
- BERT models exhibited significant errors on highly similar sentence pairs with negation or altered word order.
- BERT models were substantially slower (20-50x) than CNN and BioSentVec, impacting real-time application feasibility.
Conclusions:
- While DL models achieve high STS accuracy, their effectiveness and efficiency must be critically evaluated.
- Robustness and inference speed are key considerations for deploying STS models in clinical settings.
- Further research should focus on generalization, user-level testing, and developing diverse clinical STS datasets.
More Related Videos
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025