Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference
Thomas Savage1, Stephen P Ma2, Abdessalem Boukil3
1Division of Hospital Medicine, Perelman School of Medicine, Department of Medicine, University of Pennsylvania, 3400 Spruce St, Philadelphia, PA, 19147, United States, 1 2155191670.
Direct preference optimization (DPO) fine-tuning enhances large language model (LLM) performance on complex medical tasks like clinical reasoning and summarization, while supervised fine-tuning (SFT) suffices for simpler classification tasks.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) show potential in medicine but often require fine-tuning for optimal performance.
- Supervised fine-tuning (SFT) and direct preference optimization (DPO) are key LLM fine-tuning techniques.
- Guidance on selecting between SFT and DPO in clinical settings is limited.
Purpose of the Study:
- To compare the effectiveness of SFT and DPO for various medical natural language tasks.
- To provide data-driven recommendations for clinical informaticists on choosing fine-tuning methods.
- To inform the development and deployment of LLMs in healthcare operations.
Main Methods:
- Utilized Llama3 8B and Mistral 7B v2 models for comparison.
- Evaluated SFT and DPO performance on four core medical NLP tasks: text classification, clinical reasoning, summarization, and triage.
- Quantified performance using accuracy, Likert scales, and F1-scores, alongside computational resource analysis.
Main Results:
- DPO significantly improved clinical reasoning (accuracy from 22% to 40%) and summarization (Likert scale from 3.93 to 4.08) compared to SFT.
- SFT was sufficient for text classification (F1-score up to 0.98), while DPO showed mixed results.
- DPO fine-tuning demanded 2-3 times more computational resources than SFT alone.
Conclusions:
- SFT is adequate for straightforward tasks like text classification.
- DPO, applied after SFT, enhances performance on complex tasks including triage and clinical reasoning.
- Findings guide the strategic deployment of LLM fine-tuning in medical applications.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Dose-Response Relationship: Selectivity and Specificity
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
