Large Language Model-Based Clinical Decision Support for Antibiotic Selection and Dose Recommendation in Hospitalized
Yang Zhang1,2, Li Li3, Chunting Tan4
1School of Biomedical Engineering, Capital Medical University, No. 10 Xitoutiao, You'anmenwai, Fengtai District, Beijing, 100069, China, 86 010-83911542.
Background:
Pneumonia is a common infectious disease, and antibiotic treatment in hospitalized patients must balance efficacy, safety, and resistance risk. However, antibiotic selection and dose adjustment still rely heavily on clinician experience. Although large language models (LLMs) are promising for clinical reasoning, their direct use for antibiotic selection and dose recommendation is limited by hallucinations and weak adherence to clinical constraints.
Objective:
This study aimed to develop and externally validate a constrained LLM-based clinical decision support pipeline for antibiotic selection and dose recommendation in hospitalized patients with pneumonia.
Methods:
We conducted a multicenter retrospective study using electronic health record narratives, antibiotic orders, and laboratory indicators of hepatic and renal function from 331 hospitalized patients with pneumonia from 2 hospitals in China. The development cohort included 233 patients, and the external validation cohort included 98 patients. The pipeline integrated dual-branch retrieval (similar-case vector retrieval plus guideline-based knowledge graph retrieval), clinician-defined rule constraints, and hybrid-context reasoning. DeepSeek-V3, GLM-4.6, and GPT-4o were evaluated using F1-score and Jaccard accuracy.
Results:
On the internal test set, the full pipeline using DeepSeek-V3 achieved the best performance, with an F1-score of 0.8110 (95% CI 0.7371-0.8762) and Jaccard accuracy of 0.7624 (95% CI 0.6810-0.8386) for antibiotic selection and an F1-score of 0.7538 (95% CI 0.6671-0.8329) and Jaccard accuracy of 0.7076 (95% CI 0.6145-0.7938) for joint antibiotic selection plus dosing recommendation. On the external validation set, performance remained high, with an F1-score of 0.8605 (95% CI 0.7891-0.9252) and Jaccard accuracy of 0.8571 (95% CI 0.7857-0.9184) for antibiotic selection, and an F1-score of 0.8503 (95% CI 0.7789-0.9150) and Jaccard accuracy of 0.8469 (95% CI 0.7755-0.9133) for antibiotic selection plus dosing recommendation. The system also provided traceable evidence and rule trigger information to support clinician review.
Conclusions:
A constrained, retrieval-augmented LLM pipeline improved the consistency and interpretability of antibiotic selection and dose recommendation for hospitalized patients with pneumonia and provided preliminary evidence of cross-site generalizability.
Insights
This study developed a constrained large language model (LLM) pipeline for pneumonia antibiotic selection and dosing, improving consistency and interpretability. The system demonstrated high performance and generalizability in external validation, aiding clinical decision-making.
Area of Science:
- Medical Informatics
- Computational Medicine
- Pharmacology
Background:
- Pneumonia treatment requires balancing antibiotic efficacy, safety, and resistance, often relying on clinician experience.
- Current large language models (LLMs) face challenges like hallucinations and constraint adherence for clinical decision support.
- There's a need for reliable AI tools to guide antibiotic selection and dosing in hospitalized pneumonia patients.
Purpose of the Study:
- To develop and externally validate a constrained LLM-based clinical decision support pipeline.
- To enhance antibiotic selection and dose recommendations for hospitalized pneumonia patients.
- To improve the reliability and interpretability of AI-driven antibiotic guidance.
Main Methods:
- A multicenter retrospective study involving 331 hospitalized pneumonia patients from two Chinese hospitals.
- Development of a pipeline integrating dual-branch retrieval (case-based and knowledge graph) and clinician-defined rules.
- Evaluation of LLMs (DeepSeek-V3, GLM-4.6, GPT-4o) using F1-score and Jaccard accuracy on internal and external validation sets.
Main Results:
- The full pipeline with DeepSeek-V3 achieved high F1-scores (0.8110 internal, 0.8605 external) for antibiotic selection.
- Joint antibiotic selection and dosing recommendations also showed strong performance (0.7538 internal, 0.8503 external F1-scores).
- The system provided traceable evidence and rule triggers, enhancing transparency for clinicians.
Conclusions:
- A constrained, retrieval-augmented LLM pipeline significantly improved consistency and interpretability in antibiotic recommendations.
- The system demonstrated preliminary cross-site generalizability, suggesting potential for wider clinical application.
- This approach offers a promising advancement for AI-assisted antibiotic stewardship in pneumonia care.
Related Concept Videos
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Pneumonia IV: Management
Bacterial Pneumonia Treatment
For bacterial pneumonia, antibiotics serve as the cornerstone of therapy. Initial treatment often begins with empirical antibiotics, tailored to the anticipated causative organism and adjusted based on culture results. Key antibiotic choices include:
Pneumonia III: Complications and Assessment
Acute Pyelonephritis II: Diagnostic Studies and Management
Antibiotic Selection
Clinical Significance of Antibiotic Resistance
