Related Experiment Video
Updated: Jul 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reframing ontology fact acquisition during large language model fine-tuning as a time-to-event process
Daniel B Hier1, Tayo Obafemi-Ajayi2
1Department of Neurology and Rehabilitation, University of Illinois at Chicago, Chicago, IL, United States.
Fine-tuning large language models on Human Phenotype Ontology (HPO) and Gene Ontology (GO) facts significantly improved biomedical fact retrieval. Pretraining knowledge influences how efficiently models acquire and retain these facts during training.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence
- Computational Biology
Background:
- Large language models (LLMs) often struggle with reliable retrieval of specific biomedical facts post-pretraining.
- Fine-tuning LLMs is a strategy to enhance their ability to access and utilize specialized knowledge.
- Ontology term-identifier mappings offer structured data for evaluating knowledge acquisition in LLMs.
Purpose of the Study:
- To investigate how fine-tuning affects the acquisition of biomedical facts within LLMs.
- To analyze the role of pretraining knowledge in the fine-tuning process for fact retrieval.
- To establish a time-to-event framework for assessing fact acquisition and loss during LLM training.
Main Methods:
- LLMs were fine-tuned using Human Phenotype Ontology (HPO) and Gene Ontology (GO) facts.
- Fact acquisition was modeled as a time-to-event process indexed by training epochs.
- Stochastic decoding was employed to identify latent knowledge, distinguishing between trained and untrained fact acquisition and fact loss.
Main Results:
- Supervised fine-tuning substantially increased the correct retrieval of HPO and GO facts.
- Latent knowledge identified via stochastic decoding positively influenced the efficiency of fact acquisition.
- Untrained fact acquisition was less common, and fact loss was more frequent for untrained facts, indicating the importance of continued training exposure.
Conclusions:
- Ontology facts serve as a valuable model for studying LLM fact acquisition during fine-tuning.
- A time-to-event framework effectively characterizes the dynamics of fact acquisition, including timing and retention.
- Pretraining-derived latent knowledge is crucial for the rate and stability of fact acquisition in fine-tuned LLMs.
Related Concept Videos
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Framing Effects
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Components of Language
Automatic Processing and Automatic Social Behavior