Related Experiment Video
Updated: Sep 13, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
684
Large Language Model Symptom Identification From Clinical Text: Multicenter Study
Andrew J McMurry1,2, Dylan Phelan1, Brian E Dixon3,4
1Computational Health Informatics Program, Boston Children's Hospital, 401 Park Drive, LM5506, Mail Stop BCH3187, Boston, MA, 02215, United States, 1 617-355-4145.
Journal of Medical Internet Research
|July 31, 2025
Summary
Large language models (LLMs) accurately identify infectious respiratory disease symptoms in electronic health records, outperforming traditional methods. GPT-4 demonstrated superior accuracy and generalizability across multiple healthcare settings.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Natural Language Processing
Background:
- Patient symptom recognition is crucial for medicine, research, and public health.
- Symptoms are often underreported in coded data despite being documented in physician notes.
- Large language models (LLMs) show potential for identifying symptoms by mimicking human chart reviewers.
Purpose of the Study:
- To measure the accuracy of LLMs in identifying infectious respiratory disease symptoms based on chart review guidelines.
- To evaluate the generalizability of LLMs across different healthcare sites without site-specific adjustments.
Main Methods:
- Four LLMs (GPT-4, GPT-3.5, Llama2 70B, Mixtral 8×7B) were prompted to act as chart reviewers.
- LLM performance was optimized using a development corpus and tested against expert-annotated ground truth.
- LLM generalizability was assessed using a validation corpus from multiple emergency departments.
Main Results:
- All tested LLMs significantly outperformed the International Classification of Diseases, Tenth Revision (ICD-10)-based method (F1-score=45.1%).
- GPT-4 achieved the highest accuracy (F1-score=91.4%) and demonstrated superior generalizability in the validation corpus (F1-score=94.0%), outperforming the ICD-10 method (F1-score=26.9%).
- LLMs showed high accuracy in identifying symptoms, with GPT-4 significantly outperforming other models and the baseline method.
Conclusions:
- LLMs significantly enhance respiratory symptom identification in electronic health records compared to ICD-10 methods.
- GPT-4 exhibits high accuracy and generalizability, suggesting potential for augmenting or replacing traditional symptom identification approaches.
- LLMs can effectively mimic human chart reviewers for symptom identification, warranting further investigation into broader symptom types and clinical settings.
Keywords:
artificial intelligenceclinical text miningelectronic health recordsemergency medical servicesepidemiologic methodsinfectious disease surveillancelarge language modelsmedical informaticsnatural language processingsymptom recognition
