Related Experiment Video
Updated: Apr 14, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
AI-first, expert-verified: validating generative AI for HFACS-based coding of healthcare RCA transcripts with
Jiun-Yih Lee1, Chien-Hsien Huang2,3, Chung-Ho Hsieh4,5
1Center for Quality Management, Shin Kong Wu Ho-Su Memorial Hospital, Taipei, Taiwan.
Background:
Root cause analysis (RCA) is widely used in healthcare incident investigation, but its outputs can be limited by inconsistent causal framing and variable integration of human factors. The Human Factors Analysis and Classification System (HFACS) offers a structured taxonomy for causal attribution, yet manual coding is resource-intensive. Empirical validation of generative artificial intelligence (AI) for document-level HFACS coding from complete healthcare RCA transcripts remains limited.
Methods:
We conducted a cross-sectional validation study at an 829-bed medical center in Taiwan. Thirty-five de-identified RCA interview transcripts (2024-5) with verbatim transcription were analyzed using SKH-AI, an in-house platform integrating an Azure OpenAI-hosted GPT-4o model with deterministic decoding (temperature = 0; top_p = 1.0). The model processed each transcript holistically to identify salient narrative segments and assign HFACS codes with evidence-linked rationales without rule-based post-processing. Outputs were compared with dual-expert HFACS coding with adjudicated consensus. Performance was assessed using precision, recall, micro-/macro-F1, and Cohen's κ with bootstrapped 95% confidence intervals.
Results:
Across 562 AI-derived segments, micro-F1 was 0.66 [95% confidence interval (CI): 0.63-0.69] and macro-F1 was 0.68 (95% CI: 0.64-0.72), with moderate agreement versus expert coding (κ = 0.56, 95% CI: 0.52-0.60). Performance was higher for text-anchored categories (Level 2 Preconditions F1 = 0.70; Level 1 Unsafe Acts F1 = 0.69) than for more abstract domains (Level 3 F1 = 0.66; Level 4 F1 = 0.65). Subcategory analyses showed stronger detection of concrete cues (e.g. decision errors, physical environment, communication) and weaker performance for latent constructs (e.g. process management). Bias analyses indicated a recall-leaning tendency at Levels 3-4, consistent with increased over-attribution risk.
Conclusions:
Generative AI can produce auditable, evidence-linked candidate HFACS attributions from document-level RCA transcripts with moderate concordance to expert coding. Higher-level supervisory and organizational attributions remain vulnerable to overgeneralization and should be governed as decision support, with evidence-anchoring and mandatory expert sign-off for Level 3-4 codes.
Related Concept Videos
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Healthcare Agencies II
Parish nursing is a growing specialty nursing profession that focuses on holistic healthcare, health promotion, and illness prevention. It blends professional nursing practice with a health ministry, focusing on health and healing within the context of a Christian community. Parish nurses serve as health educators, referral sources,...
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Healthcare Agencies I
Heuristics
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...