Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Generating Human-Readable Labels for SNOMED CT Expressions with LLMs: A Study on Model Performance and Rater

Adel Bensahla1,2, Jamil Zaghir1,2, Yuanyuan Zheng1,2

  • 1Division of Medical Information Sciences, Geneva University Hospitals, Geneva, Switzerland.

Studies in Health Technology and Informatics
|May 23, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A gold-standard French-language annotated corpus of oncological entities with ICD-O normalisation.

Scientific data·2026
Same author

Clarifying the relationship between biomedical and health informatics and digital health: expert perspectives.

BMJ health & care informatics·2026
Same author

QT prolongation alerts lead to monitoring but rarely to therapeutic changes: a prospective hospital study.

Frontiers in pharmacology·2026
Same author

Can LLMs Turn French PET/CT Narrative Reports into Structured Knowledge?

Studies in health technology and informatics·2026
Same author

Explainable Framework for Ontology-Based Similarity: A Use Case on SNOMED CT.

Studies in health technology and informatics·2026
Same author

Advancing Knowledge in Evaluating the Clinical Impact of Large Language Models for Clinical Text Summarization: A Narrative Review.

Studies in health technology and informatics·2026

Large language models (LLMs) can generate human-readable labels for post-coordinated SNOMED CT expressions. While LLM labels are semantically equivalent to official terms, expert agreement on novel expressions is low, highlighting task subjectivity.

Area of Science:

  • Medical Informatics
  • Natural Language Processing
  • Clinical Terminology

Background:

  • Post-coordinated SNOMED CT expressions lack human-readable labels, limiting their utility in data analysis and machine learning.
  • Developing automated methods for labeling these expressions is crucial for broader adoption and application.

Purpose of the Study:

  • To evaluate the quality of labels for post-coordinated SNOMED CT expressions generated by a large language model (LLM).
  • To assess the semantic equivalence and phrasing quality of LLM-generated labels compared to official SNOMED CT terms.

Main Methods:

  • A large language model (LLM) generated labels for 100 post-coordinated SNOMED CT expressions.
  • Two clinical terminologists blindly rated the LLM-generated labels for semantic equivalence and phrasing quality.
Keywords:
Large Language Model (LLM)Representation LearningSNOMED CTpost-coordination

Related Experiment Videos

Last Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Main Results:

  • LLM-generated labels demonstrated semantic equivalence comparable to official SNOMED CT terms.
  • Errors in LLM labels were primarily stylistic, not conceptual.
  • Interrater reliability between experts was negligible when rating novel post-coordinated expressions.

Conclusions:

  • Automated labeling using LLMs is a viable approach for SNOMED CT expressions.
  • The main challenge lies in the inherent subjectivity of labeling novel expressions, not in model performance.
  • The goal should shift towards achieving semantic plausibility at scale rather than a single expert consensus.