Related Experiment Video
Updated: Sep 18, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Ontology enrichment using a large language model: Applying lexical, semantic, and knowledge network-based similarity
Navya Martin Kollapally1, James Geller2, Vipina Kuttichi Keloth3
1Kean University, United States.
This study developed an automated pipeline to enrich ontologies, specifically the Social Determinants of Health Ontology (SDoH), using Large Language Models (LLMs) and PubMed. The LLM-based approach successfully identified more relevant concepts than existing databases, enhancing the ontology.
Area of Science:
- Biomedical Informatics
- Knowledge Representation
- Ontology Engineering
Background:
- Ontologies are crucial for domain knowledge representation and require comprehensive domain views for utility.
- Ontology enrichment is necessary to incorporate new concepts and advancements in the field.
- The Social Determinants of Health (SDoH) domain requires continuous updates to its knowledge representation.
Purpose of the Study:
- To develop an automatic pipeline for ontology enrichment using a seed ontology, a Large Language Model (LLM), and a text corpus.
- To apply and demonstrate the effectiveness of this pipeline for extending the SDoH Ontology (SOHOv1).
- To establish a generalizable methodology for ontology enrichment in other domains.
Main Methods:
- Retrieved PubMed abstracts using existing SOHOv1 concepts as search terms.
- Employed GPT-4-1201 to extract semantic triples from abstracts, followed by filtering using lexical, semantic, and knowledge network-based approaches.
- Compared the granularity of extracted triples with SemMedDB and validated results using human experts and ontology tools.
Main Results:
- Expanded SOHOv1 (173 concepts, 585 axioms) to SOHOv2 (572 concepts, 1,542 axioms), significantly increasing concept and axiom counts.
- The LLM-based method identified a greater number of concepts compared to those extracted from SemMedDB.
- Demonstrated the feasibility and effectiveness of the automated enrichment pipeline for the SDoH domain.
Conclusions:
- Successfully extracted semantic triples from PubMed abstracts using GPT-4-1201 via prompt chaining.
- Demonstrated the superiority of GPT-4-1201 extracted triples over SemMedDB for SDoH ontology enrichment.
- Utilized advanced search techniques for concept identification and confirmed concept quality through expert evaluation, validating the generalizability of the methodology.
More Related Videos
Related Concept Videos
Natural and Artificial Concepts
Concepts and Prototypes
The brain organizes this information using concepts, which are mental categories grouping linguistic data,...
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Causes of Similarity-Dissimilarity Effect

