Related Experiment Video
Updated: Aug 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking MeSH-augmented embeddings for biomedical document similarity
Rohitha Ravinder1,2, Lukas Geist3, Nelson Quiñones3
1ZB MED Information Centre for Life Sciences, Gleueler Str. 60, 50931, Cologne, Germany. ravinder@zbmed.de.
Journal of Biomedical Semantics
|August 14, 2026
Summary
Integrating Medical Subject Headings (MeSH) with document embeddings improves biomedical literature retrieval. Hybrid methods show benefits, with fine-tuned transformer models like BioBERT achieving the highest performance.
Area of Science:
- Biomedical Informatics
- Information Retrieval
- Natural Language Processing
Background:
- Efficient retrieval of biomedical literature is crucial due to its vast volume.
- Embedding-based methods enhance document retrieval over traditional keyword approaches.
- Integrating domain-specific terminologies like Medical Subject Headings (MeSH) into embedding models is underexplored.
Purpose of the Study:
- To compare hybrid methods integrating MeSH annotations with document embeddings.
- To benchmark these hybrid methods against traditional and transformer-based models.
- To evaluate the impact of MeSH integration on biomedical document retrieval performance.
Main Methods:
- Three hybrid methods (pre-annotation, post-annotation, post-reduction) integrating MeSH with document embeddings were developed.
- Methods were benchmarked against TF-IDF, Word2Vec, fastText, Doc2Vec, BioBERT, SciBERT, SPECTER, and SapBERT.
- Experiments utilized the RELISH corpus with cosine similarity and Word Mover's Distance (WMD) for evaluation.
Main Results:
- Fine-tuned BioBERT and SciBERT models demonstrated the best alignment with expert judgments.
- Doc2Vec and MeSH-based hybrid methods performed well, highlighting the value of combining vocabularies with embeddings.
- While MeSH integration offered modest improvements (2-4%), fine-tuned large models achieved up to 90% precision.
Conclusions:
- Performance gains from concept integration may be moderate, but the study successfully benchmarks structured vocabularies with embedding methods.
- These techniques are applicable for aligning literature with other data sources using controlled vocabularies.
- The Dockerized pipeline ensures reproducibility and supports future research in biomedical document retrieval.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Bioequivalence: Overview
Pharmaceutical equivalents, by definition, are drug products with the same active ingredient in the same quantities, encapsulated in identical dosage forms, and intended for the same administration routes. These pharmaceutical equivalents are deemed bioequivalent if the bioavailability of the active entity in the drug preparations is similar. Moreover, pharmaceutical equivalents demonstrating bioequivalence are also regarded as therapeutically equivalent. This means that when used as directed,...
