Related Experiment Video
Updated: Jun 26, 2026

A Method for Manipulating Blood Glucose and Measuring Resulting Changes in Cognitive Accessibility of Target Stimuli
Published on: August 12, 2016
Fine-tuning and structured prompting strategies for question answering over full-text biomedical research articles
Kaiming Tao1, Rohit Satija1, Jinru Zhou1
1Division of Infectious Diseases, Dept. of Medicine, Stanford University, Stanford, California, United States of America.
Objectives:
The ability of large language models (LLMs) to answer targeted scientific questions by synthesizing information from research articles remains an open research challenge.
Methods:
We evaluated the effects of fine-tuning and a question-specific prompting strategy to answer 16 pre-defined questions about HIV drug resistance studies, including whether viral genetic sequences were reported and the demographics and antiviral treatments of the individuals from whom sequences were obtained. For fine-tuning, we constructed an instruction set comprising 250 HIV drug resistance studies, with 16 questions per study and corresponding answers and explanations. For question-specific prompting, we developed a set of if-then rules tailored to each question. We compared the performance of three base models - GPT-4o-mini-2024-07-18 (GPT-4o), Meta Llama-3.1-70B-Instruct (Llama-3.1-70B), and Meta Llama-3.1-8B-Instruct (Llama-3.1-8B) - with their performance using fine-tuning, question specific prompting, and fine-tuning followed by question-specific prompting. Performance was assessed using accuracy, precision, recall, and F1 score, averaged over 150 held-out studies not used for fine-tuning. Comparisons were performed using Wilcoxon signed-rank tests.
Results:
Fine-tuning increased precision by 5% for GPT-4o, 16% for Llama-3.1-70B, and 8% for Llama-3.1-8B, although this increase reached statistical significance only for Llama-3.1-70B. Fine-tuning also significantly increased recall for GPT-4o by 11%. Question specific prompting increased recall for all three models (6% for GPT-4o, 7% for Llama-3.1-70B, and 18% for Llama-3.1-8B), with statistically significant improvements observed only for Llama-3.1-8B. Applying question specific prompting to each of the fine-tuned models did not yield additional improvements beyond fine-tuning alone. When pooled across the three models, fine-tuning was associated with a greater effect on precision than recall (OR = 4.35; p = 0.001; Fisher's exact test), whereas question-specific prompting led to a greater effect on recall than on precision (OR= 7.09; p = 0.0001; Fisher's exact test).
Conclusions:
In this domain-focused proof-of-concept study, fine-tuning and question-specific prompting each led to improvement in one or more metrics for each of the three models. Pooled analyses indicated that fine-tuning improved precision, whereas question specific prompting preferentially improved recall.
Related Concept Videos
Techniques of Therapeutic Communication II: Focusing, Paraphrasing, and Summarizing
This therapeutic technique can also be used when a patient brings up pertinent information during a health-related conversation. The...
Catalytically Perfect Enzymes