Related Experiment Video
Updated: Jul 4, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
561
Comparing the Quality of Domain-Specific Versus General Language Models for Artificial Intelligence-Generated
Alireza Akhondi-Asl1,2,3,4, Youyang Yang1,2,3,4, Matthew Luchette1,2,3,4
1Division of Critical Care Medicine, Department of Anesthesiology, Critical Care and Pain Medicine, Boston Children's Hospital, Boston, MA.
Summary
A smaller, domain-adapted language model (LM) fine-tuned on pediatric critical care notes outperformed larger general LMs in generating differential diagnoses. While still inferior to clinicians, these specialized LMs show potential as adjunct tools.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
Background:
- Generative language models (LMs) are increasingly evaluated for healthcare applications, yet studies in pediatric critical care are limited.
- Assessing the utility of LMs for clinical tasks in the pediatric intensive care unit (PICU) is crucial.
Purpose of the Study:
- To evaluate the effectiveness of generative LMs in the PICU setting.
- To compare the performance of domain-adapted LMs against larger, general-domain LMs in generating differential diagnoses from patient admission notes.
Main Methods:
- Retrospective cohort study utilizing admission notes from a quaternary PICU (January 2012-April 2023).
- Development and validation using over 1.9 million notes from 32,454 patients.
- Evaluation of differential diagnoses generated by clinicians and various LMs (general and fine-tuned) by five critical care experts using a 5-point Likert scale.
Main Results:
- The best-performing model, a fine-tuned LLaMa-7B, achieved a mean quality score of 2.88, compared to clinicians' score of 3.43.
- Fine-tuned LLaMa-7B significantly outperformed larger general LMs, including LLaMa-65B and BioGPT-Large.
- Clinicians' diagnoses were ranked highest quality in 55% of cases, while fine-tuned LLaMa-7B was ranked highest in 29%.
Conclusions:
- A smaller language model fine-tuned on specific domain data (PICU notes) can outperform larger, general-domain models.
- Current LMs are not superior to human clinicians but show promise as supplementary tools for real-world clinical tasks.

