Related Experiment Video
Updated: Sep 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
LLM-based approaches for automated vocabulary mapping between SIGTAP and OMOP CDM concepts
Vinícius João de Barros Vanzin1, Dilvan de Abreu Moreira1, Ricardo Marcondes Marcacini1
1Institute of Mathematics and Computer Sciences (ICMC) - University of Sao Paulo (USP), Av. Trab. São Carlense, 400, São Carlos, 13566-590, SP, Brazil.
Mapping Brazilian healthcare terminologies (SIGTAP) to OMOP Common Data Model (CDM) terminologies using Large Language Models (LLMs) significantly reduces expert effort. LLM-based approaches demonstrate effective vocabulary mapping for medicines and medical procedures.
Area of Science:
- Health Informatics
- Medical Terminology
- Artificial Intelligence in Healthcare
Background:
- Global healthcare systems face challenges in integrating diverse medical terminologies and classification systems.
- The widespread adoption of Electronic Health Record (EHR) systems necessitates standardized information exchange.
- Bridging the gap between national terminologies like Brazil's SIGTAP and international standards such as OMOP CDM is crucial.
Purpose of the Study:
- To develop and evaluate methods for mapping the Brazilian SIGTAP vocabulary to OMOP Common Data Model (CDM) terminologies.
- To assess the effectiveness of two distinct Large Language Model (LLM)-based pipelines for vocabulary mapping.
- To compare the performance of LLM pipelines in mapping SIGTAP medicines and medical procedures.
Main Methods:
- Two pipelines were developed for vocabulary mapping: one using textual embeddings and retrieval-augmented generation (RAG) with LLMs, and another using LLM agents with predefined protocols.
- The pipelines were evaluated on subsets of the SIGTAP vocabulary, specifically medicines and medical procedures.
- Performance metrics, including F1-score and recall, were used to compare the pipelines' effectiveness.
Main Results:
- Both LLM-based pipelines achieved comparable performance in mapping procedures (F1: 0.684 vs. 0.678) and medicines (F1: 0.846 vs. 0.839).
- The second pipeline, utilizing LLM agents with dynamic query refinement, showed an advantage in recall.
- LLM-based methods substantially decreased the manual effort required from domain experts.
Conclusions:
- LLM-based approaches are viable and effective for mapping national healthcare terminologies to standardized models like OMOP CDM.
- The agent-based pipeline with dynamic query refinement offers improved recall, enhancing mapping accuracy.
- These findings support the use of AI in healthcare informatics to streamline terminology integration and reduce expert workload.
More Related Videos
12:26Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
Published on: October 11, 2016
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018