Related Experiment Video
Updated: Jun 12, 2025

Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
Published on: November 1, 2024
The reliability of deep learning models in assessing the shoulder arthroscopic field's visual clarity in relation to
Son Quang Tran1,2, Minh Cong Bui1,2, Dat Tien Nguyen3
1Clinical Sciences Program, Faculty of Medicine, Chulalongkorn University, Bangkok, Thailand.
Background:
The clarity of visualization in shoulder arthroscopy is significantly influenced by intraoperative bleeding. This study aims to develop deep learning models to classify the visual clarity of arthroscopic shoulder images and evaluate their reliability compared to rater assessment.
Methods:
We retrospectively reviewed videos from 113 shoulder arthroscopies, using 63 patients' videos to create a 3750-image training dataset and 50 patients' videos to evaluate the reliability and agreement of the trained models. Images extracted from the videos were assessed for visual clarity with a 3-grade scale. Subsequently, we implemented transfer learning techniques for the pretrained deep learning models involving DensetNet169, DenseNet201, Xception, InceptionResNetV2, VGG16, and ViT. The reliability and agreement of the trained predictive models compared with raters in classifying the visual clarity of shoulder arthroscopic images were reported with percent agreement and weighted kappa coefficients. Similarly, for quantifying the visual clarity of surgical videos, the reliability and agreement were evaluated through intraclass coefficients, bias, and limits of agreement.
Results:
Most models achieved over 90% accuracy in validation step, with InceptionResNetV2, DenseNet169, and ViT exhibiting good percent agreements of 82.3%, 79.2%, and 77.8%, respectively, and their weighted kappa coefficients above 0.8. DenseNet169 model had the highest reliability for evaluating the visual clarity of surgical videos with an intraclass correlation coefficient of 0.9, a bias of 0.027, and the narrowest limits of agreement.
Conclusion:
The trained DenseNet169 predictive model was reliable enough to be utilized as an objective measure of the visual clarity of the shoulder arthroscopic field for further research.

