Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.6K
3.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

"Check Symptoms & Get Care": Mount Sinai's AI Triage Solution.

NEJM catalyst innovations in care delivery·2026
Same author

Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.

Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology·2026
Same author

Rare protein-coding variation and the genetic architecture of height in >1.4 million individuals.

medRxiv : the preprint server for health sciences·2026
Same author

Lewy pathology largely absent in prefrontal cortices of Parkinson's disease patients undergoing deep brain stimulation.

NPJ Parkinson's disease·2026
Same author

Regional, functional and transcriptomic decoding of multidimensional brain structure alterations in obsessive-compulsive disorder.

Nature communications·2026
Same author

Evaluating Sycophancy in Frontier Models Using Persona-Driven Challenge.

medRxiv : the preprint server for health sciences·2026

Related Experiment Video

Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

A scalable framework for benchmark embedding models in semantic health-care tasks.

Shelly Soffer1,2, Mahmud Omar3,4, Moran Gendler5

  • 1Institute of Hematology, Davidoff Cancer Center, Rabin Medical Center, Petah Tikva, 49100, Israel.

Journal of the American Medical Informatics Association : JAMIA
|September 22, 2025
PubMed
Summary

A new benchmarking method evaluates text embeddings for healthcare semantic tasks. Larger models excel in long-context retrieval, while smaller models perform comparably in short tasks.

Keywords:
artificial intelligencelarge language modelsmedical NLP tools

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K

Related Experiment Videos

Last Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K

Area of Science:

  • Natural Language Processing
  • Biomedical Informatics
  • Machine Learning

Background:

  • Text embeddings are crucial for semantic tasks like retrieval augmented generation (RAG).
  • Healthcare applications of text embeddings are limited by a lack of standardized benchmarking.
  • Existing methods are insufficient for evaluating embedding models in complex healthcare scenarios.

Purpose of the Study:

  • To introduce a scalable benchmarking method for evaluating text embedding models in healthcare.
  • To assess the performance of various embedding models across diverse medical semantic tasks.
  • To provide a framework for selecting optimal embedding models for healthcare applications.

Main Methods:

  • Evaluated 39 embedding models on 7 medical semantic similarity tasks.
  • Utilized diverse datasets including patient data (Mount Sinai, MIMIC IV), PubMed articles, and synthetic data.
  • Assessed semantic textual similarity (STS) using Spearman rank correlation and reframed tasks as retrieval problems evaluated by mean reciprocal rank and recall at k.

Main Results:

  • Larger embedding models (>7b parameters) generally outperformed smaller models, especially in long-context tasks.
  • NV-Embed-v1 (7b) excelled in short tasks but underperformed in long-context scenarios.
  • Smaller models (e.g., b1ade-embed, 335M) matched larger models in short tasks, while larger models significantly led in long retrieval tasks.

Conclusions:

  • The developed benchmarking framework is scalable and flexible for healthcare semantic tasks.
  • This structured approach guides the selection of appropriate embedding models for diverse healthcare applications.
  • Effective model selection enhances critical applications such as semantic search and retrieval-augmented generation (RAG).