Related Experiment Video
Updated: Apr 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Investigating fine-tuning versus zero-shot learning for general large language models when predicting cancer survival
T Phaterpekar1, Z Zeng2, Y Mali3
1Faculty of Medicine, University of British Columbia, Vancouver, Canada.
Background:
Unstructured oncology consultation notes contain rich clinical information that may support survival prediction. Open-weight large language models (LLMs) can utilize these notes with zero-shot inference or fine-tuning, but their relative value for this setting remains unclear. The objective of this study is to evaluate open-weight LLMs for predicting 60-month survival from initial oncology consultation notes, comparing (i) zero-shot performance, (ii) performance after fine-tuning, and (iii) smaller natural language processing models trained on the same dataset in prior work.
Materials And Methods:
We used Meta's Llama models to predict patients' 60-month survival using oncology consultation notes from a dataset of 59 800 patients. We tested both zero-shot and fine-tuning approaches. Metrics included balanced accuracy (BA) and weighted F1.
Results:
Zero-shot performance was limited. Llama-2-13B performed best among the zero-shot configurations (average performance across prompts: BA 0.596, weighted F1 0.644; performance on Prompt 4: BA 0.766, weighted F1 0.802). Fine-tuning improved performance across models: Llama-2-13B achieved BA 0.842, weighted F1 0.846, area under the receiver operating characteristic curve (AUC) 0.905; Llama-2-7B achieved BA 0.840, weighted F1 0.843, AUC 0.911; Llama-3.1-8B achieved BA 0.829, weighted F1 0.829, AUC 0.881. Performance was numerically similar to smaller models trained on the same task and data.
Conclusions:
For predicting 60-month survival from initial oncology consultation documents, fine-tuning open-weight LLMs meaningfully improves performance compared with zero-shot use, but does not consistently outperform smaller language models. This may suggest that both fine-tuned LLMs and smaller models merit continued investigation, with the most appropriate approach likely to depend on the outcome of interest, clinical context, and practical considerations such as hardware, privacy, and deployment feasibility.
Related Concept Videos
Cancer Survival Analysis
Improving Translational Accuracy
Improving Translational Accuracy
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Targeted Cancer Therapies
There are several types of targeted therapies against...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
