Related Experiment Videos
Robust and Interpretable AI for Acute Appendicitis: A Simulation-to-Clinical Validation Pipeline
Kazi Nur Uddin1, Ebrima Njie1, Ruoming Jin2
1Department of Mathematical Sciences, Kent State University, Kent, OH, 44242, United States.
Background And Objective:
Accurate and timely diagnosis of Acute Appendicitis remains a major challenge in emergency medicine, largely because of its overlapping symptoms and heterogeneous clinical presentations. This study introduces the Robust and Interpretable Simulation-Analysis (RISA) framework, a dual-phase methodology designed to evaluate diagnostic classifiers for robustness, generalizability, and interpretability.
Methods:
Six supervised learning algorithms (Fisher Discriminant Analysis, Linear Discriminant Analysis, Quadratic Discriminant Analysis, Logistic Regression, Support Vector Machine, and Random Forest) were systematically assessed through 24 factorial simulation scenarios that varied in class imbalance, dimensionality, covariance structure, and nonlinearity of decision boundaries. The simulation phase was followed by clinical validation on two independent datasets: the adult Appendicitis dataset (n=106) and the Regensburg Pediatric Appendicitis (RPA) dataset (n=782). Model interpretability was investigated using Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP).
Results:
In simulation experiments, Linear Discriminant Analysis and Support Vector Machine exhibited balanced sensitivity and specificity, whereas Random Forest displayed the highest resilience under nonlinear and heterogeneous situations. In clinical validation, Linear Discriminant Analysis attained the greatest AUC of 0.908 on the appendicitis dataset and 0.729 on the RPA dataset. For the Appendicitis dataset, Fisher Discriminant Analysis demonstrated optimal sensitivity (1.00), whereas Quadratic Discriminant Analysis attained flawless specificity (1.00). In the RPA cohort, the Support Vector Machine demonstrated the highest overall discrimination (AUC = 0.762) with consistent fold-wise performance. LIME and SHAP consistently identified clinically established biomarkers, including white blood cell count, neutrophil percentage, and body temperature, supporting the medical validity of the model findings.
Conclusions:
The RISA framework offers a reproducible and transparent methodology for creating reliable diagnostic artificial intelligence by integrating simulation-based benchmarking with interpretable clinical validation. It connects algorithmic robustness with clinical reliability, facilitating the implementation of explainable machine learning models for the diagnosis of Acute Appendicitis.
Related Concept Videos
Appendicitis-II: Diagnostic Studies and Management
Diagnosing Appendicitis
It requires a multifaceted approach, starting with a detailed physical examination to pinpoint the location and nature of the pain and identify any associated symptoms. Laboratory tests play a crucial role. A complete Blood Count (CBC) typically reveals leukocytosis (an increased number of...
Appendicitis-I: Introduction
Etiology: Appendicitis can arise from various causes, primarily rooted in the obstruction of the appendix lumen. Factors contributing to this obstruction include fecal accumulation, lymphoid hyperplasia and, in...