Related Experiment Video
Updated: Jul 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Development of a benchmarking dataset for symptom detection using large language models
Joshua Davis1,2, Brigitte N Durieux1, Chloe Van Dongen1
1Department of Supportive Oncology, Dana-Farber Cancer Institute, Boston, MA, 02215, United States.
Objectives:
To develop a pipeline for evaluating large language models (LLMs) on the task of capturing symptoms from clinical encounters.
Materials And Methods:
We created a gold standard dataset of symptom annotations from simulated doctor-patient encounter excerpts (264 encounters; 16 symptoms; double-coded and adjudicated). Nine different LLMs from 4 vendors (OpenAI, Meta, DeepSeek, Moonshot AI) were used as examples to test our evaluation pipeline; outputs were assessed for correct structure and symptom information.
Results:
Of 3085 excerpts, 2087 (68%) contained symptoms. Pain, cough, and shortness of breath were most common; LLMs achieved F1 scores ranging 0.66-0.88 for these symptoms with minimal prompt engineering. Of tested models, GPT-4.1 demonstrated the best overall performance.
Discussion:
Our evaluation pipeline and benchmarking dataset are publicly available and applicable to various LLMs, including open-source models.
Conclusion:
This work supports the development and optimization of models that seek to improve patient symptom understanding.