Related Experiment Video
Updated: Apr 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Hidden failure modes of large language models in healthcare-associated infection surveillance: a structured
Mamdooh Alzyood1, Alfred Veldhuis1, Hayley Stevenson2
1Faculty of Health Science, and Technology, School of Psychology, Social Work and Public Health, https://ror.org/04v2twj65Oxford Brookes University Faculty of Health and Life Sciences, Oxford, UK.
Background:
Large language models (LLMs) are increasingly explored for healthcare-associated infection (HAI) surveillance, but their reliability in applying formal National Healthcare Safety Network (NHSN) definitions is not well characterized. This study evaluates GPT-5.1 Thinking's accuracy and rationales in classifying NHSN-defined infections.
Methods:
Seventy synthesized case vignettes containing complete, organized clinical data representing five NHSN infection types, including complex edge cases, were assessed using 2025 NHSN surveillance definitions. GPT-5.1 Thinking classified cases under three prompting strategies: standard, structured, and constrained. Quantitative accuracy metrics and qualitative inductive content analysis of rationales and failure modes were performed.
Results:
Overall accuracy across 210 classifications improved from 78.6% (standard prompt) to 88.6% (structured) and 95.7% (constrained). Performance was highest for infections with clear anatomical or radiographic criteria (surgical site infections [SSI], ventilator-associated pneumonia [VAP]) and lowest for infections involving complex exclusion rules (central line-associated bloodstream infection [CLABSI], Clostridioides difficile infection [CDI]). Constrained prompting enhanced adherence to NHSN rules but did not eliminate errors in hierarchical exclusions. Content analysis identified three recurrent failure categories: prioritization of clinical plausibility over surveillance logic, failure to apply quantitative and temporal thresholds, and errors in hierarchical source attribution.
Conclusion:
GPT-5.1 Thinking shows potential to support infection surveillance under strict constraints but exhibits systematic limitations, including overreliance on clinical intuition and difficulty with complex exclusion pathways. Currently, LLMs are unsuitable for autonomous NHSN classification but may serve as supervised decision-support tools with robust human oversight. Further development is needed to enhance LLMs' ability to synthesize surveillance definitions and complex situational characteristics critical for effective HAI surveillance, though fully autonomous deployment would require further validation. These findings are based on synthetic data that may differ from real-world clinical data in ways likely to overestimate the accuracy of these tools.
Related Concept Videos
Healthcare Associated Infections II: Preventive Measures
The best practices for preventing healthcare-associated infections include hand hygiene, patient risk...
Healthcare Associated Infections I: Iatrogenic, Exogenic and Endogenic
HAIs significantly increase the cost of health care. Extended stays in healthcare institutions, increased disability, increased costs of medications, including specialized antibiotics, and prolonged recovery times add to the patient's expenses and the healthcare institution and funding bodies.
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Principles of Disease Surveillance
Steps in Outbreak Investigation
Factors Affecting the Risk of Infection
The integrity and count of the white blood cells help the body resist pathogens and fight infection. When impaired, it reduces the body's resistance to pathogens. The acidic pH levels of the gastrointestinal, genitourinary tracts, and skin...