Related Experiment Video
Updated: Sep 16, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
692
Benchmarking vision-language models for diagnostics in emergency and critical care settings
Christoph F Kurz1, Tatiana Merzhevich2, Bjoern M Eskofier2,3
1Novartis Pharma GmbH, Nuremberg, Germany. christoph.kurz@novartis.com.
NPJ Digital Medicine
|July 10, 2025
Summary
Vision-language models (VLMs) show limited diagnostic accuracy in acute care settings. Specialized training is needed to improve open-source VLMs for emergency and intensive care applications, as GPT-4o demonstrated superior performance.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Clinical Decision Support Systems
Background:
- Vision-language models (VLMs) are emerging AI tools with potential in healthcare.
- Their application in high-acuity clinical settings like emergency and intensive care units is not well understood.
- Current open-source VLMs may not meet the diagnostic demands of acute care.
Purpose of the Study:
- To evaluate the diagnostic performance of small open-source VLMs in acute care scenarios.
- To compare the capabilities of open-source VLMs against a state-of-the-art model (GPT-4o).
- To identify limitations and areas for improvement in VLMs for emergency medicine.
Main Methods:
- Utilized a multimodal dataset comprising diagnostic questions with medical images and clinical context.
- Benchmarked several small, open-source vision-language models.
- Compared performance against the advanced GPT-4o model.
Main Results:
- Open-source VLMs achieved a maximum diagnostic accuracy of 40.4%.
- GPT-4o significantly outperformed open-source models, reaching 68.1% accuracy.
- A substantial performance gap exists between current open-source VLMs and advanced models in acute care.
Conclusions:
- Open-source VLMs currently exhibit insufficient diagnostic accuracy for critical acute care applications.
- Significant advancements in specialized training and model optimization are required for open-source VLMs.
- GPT-4o shows promise but further validation is needed for clinical deployment in emergency and intensive care.
Related Concept Videos
SBAR II: Application of SBAR
4.8K
SBAR is an effective communication tool used by healthcare professionals to communicate patient information accurately. SBAR stands for Situation, Background, Assessment, and Recommendation. For a better understanding, an example is given below.
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
4.8K
Cardiopulmonary Resuscitation II: ACLS Airway Management
109
Airway management is a key skill in emergency and critical care settings, as maintaining a clear airway is essential for adequate oxygenation and ventilation.Head Tilt-Chin Lift TechniqueThe head tilt-chin lift maneuver is an essential technique primarily used in patients without suspected cervical spine injuries. To perform this maneuver, one hand is placed on the patient’s forehead, and gentle pressure is applied backward to tilt the head. The fingertips of the other hand are positioned...
109

