Related Experiment Video
Updated: Jul 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reframing ontology fact acquisition during large language model fine-tuning as a time-to-event process
Daniel B Hier1, Tayo Obafemi-Ajayi2
1Department of Neurology and Rehabilitation, University of Illinois at Chicago, Chicago, IL, United States.
Introduction:
Large language models can be fine-tuned to improve retrieval of biomedical facts that are not reliably accessible after pretraining. We used ontology term-identifier mappings from the Human Phenotype Ontology (HPO) and Gene Ontology (GO) as structured biomedical facts to study fact acquisition during fine-tuning.
Methods:
Each ontology fact consisted of a term and its corresponding machine-readable identifier. We reframed fact acquisition as a discrete time-to-event process indexed by training epoch. Llama-3.1-8B Instruct was fine-tuned on HPO and GO ontology facts. Deterministic retrieval accuracy was assessed at baseline and after each fine-tuning epoch. Repeated stochastic decoding at baseline was used to probe for latent parametric support. Baseline-incorrect facts that were recovered at least once under stochastic decoding were classified as latent-knowledge-positive; facts not recovered under stochastic decoding were classified as latent-knowledge-negative. We distinguished trained-fact acquisition, defined for facts included in the fine-tuning set, from untrained-fact acquisition, defined for facts withheld from training. Fact loss was defined as the first transition from correct to incorrect retrieval among facts that were correct in the base model. Kaplan-Meier estimators were used to construct fact-acquisition and fact-loss curves over training epochs, and Cox proportional hazards models were used to identify predictors of acquisition rate.
Results:
At baseline, the model correctly retrieved 1.1% (9/800) of HPO facts and 5.6% (44/802) of GO facts. After 20 epochs of supervised fine-tuning, correct retrieval increased to 71.9% (575/800) for HPO-trained facts, 61.8% (248/401) for GO-trained facts, and 11.2% (45/401) for GO-untrained facts. Latent-knowledge-positive facts were acquired more efficiently than latent-knowledge-negative facts. Corpus-based fact support in the biomedical literature had smaller positive effects on fact acquisition. Untrained-fact acquisition for withheld GO facts was uncommon, occurring in 5.8% (22/378) of baseline-incorrect withheld facts, but was more likely for latent-knowledge-positive facts. Among GO facts that were correct at baseline, fact loss was more frequent for untrained facts than for trained facts, suggesting that continued training exposure was associated with greater retention of already-correct mappings.
Discussion:
Ontology facts provide a useful experimental model for studying fact acquisition during large language model fine-tuning. A time-to-event framework reveals not only whether facts are acquired, but also when they are acquired and whether they are later lost. The findings suggest that pretraining-derived latent knowledge, detectable through stochastic decoding, influences the rate and stability of fact acquisition during fine-tuning. This framework may help evaluate fine-tuning strategies, curriculum design, and the durability of learned biomedical facts.
Related Concept Videos
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Framing Effects
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Components of Language
Automatic Processing and Automatic Social Behavior