Related Experiment Videos
Integrating Local and Global Representation Learning for Pediatric Pneumonia Detection: A Hybrid CNN-Transformer
Ece Meltem Yalçın1, Hayriye Tanyıldız2, Serpil Aslan2
1Faculty of Medicine, Firat University, 23119 Elazig, Türkiye.
Abstract:
Background/Objectives: Pneumonia remains a leading cause of childhood morbidity and mortality worldwide. Accurate interpretation of pediatric chest radiographs is challenging because of anatomical variability, subtle radiographic findings, and inter-observer variability. This study evaluates different CNN-Transformer ensemble strategies for pediatric pneumonia detection by combining complementary local and global image representations. Methods: Experiments were conducted on the publicly available Pediatric Pneumonia Chest X-ray dataset containing 5856 radiographs. EfficientNetV2-S was used to extract local features, whereas Swin Transformer-T modeled global anatomical relationships. Soft voting, weighted voting, and stacking were evaluated under a unified training protocol. Image preprocessing, data augmentation, and Weighted Random Sampling were applied to improve robustness and address class imbalance. Performance was assessed using an independent hold-out test set and five-fold cross-validation. Grad-CAM was used to interpret model predictions. Results: Ensemble learning improved classification performance compared with individual models while revealing different trade-offs among fusion strategies. The Soft Ensemble achieved the highest hold-out accuracy (96.96%) and F1-score (97.59%). The Hybrid CNN-Transformer Stacking model achieved the highest sensitivity (99.49%) and produced the fewest false-negative predictions (n = 2), while demonstrating the most consistent performance across five-fold cross-validation. Grad-CAM visualizations indicated that the CNN and Transformer models captured complementary radiographic information. Conclusions: The proposed framework demonstrates that different ensemble strategies offer distinct advantages. Soft voting provided the best overall hold-out performance, whereas stacking minimized false-negative predictions and achieved the highest sensitivity, indicating its potential for AI-assisted pediatric pneumonia screening.