Related Experiment Video
Updated: Aug 6, 2026

Flexible Colonoscopy in Mice to Evaluate the Severity of Colitis and Colorectal Tumors Using a Validated Endoscopic Scoring System
Published on: October 16, 2013
Computer Vision Approach to Triaging Patient-Submitted Photos of Intestinal Ostomies
Chris Varghese1,2, Ashok Choudhary1, Ellen Larson1,3
1Division of Hepatobiliary and Pancreas Surgery, Mayo Clinic, Rochester, Minnesota.
Background:
Surgeons and ostomy nurses receive a high volume of stoma photos from patients.
Objective:
This study aimed to develop and validate an automated and scalable artificial intelligence pipeline for identification and triage of patient-submitted intestinal ostomy photos.
Design:
This was a retrospective derivation and validation study with stakeholder engagement to guide model design. Clinical teams reviewed photos in duplicate and categorized them into: 1) obstructed view, 2) healthy stoma, 3) suitable for conservative management, or 4) requiring in-person review. Pretrained neural networks, including MobileNetV4, ResNet50, vision transformer, and contrastive language-image pretraining vision transformer, were fine-tuned with 5-fold cross validation.
Settings:
Nine Mayo Clinic hospitals (2019-2022).
Patients:
Adult patients (age 18 years or older) undergoing surgery who submitted images within 30 days after surgery were included in the study.
Main Outcome Measures:
Model performance was evaluated with area under the receiver operating characteristic curve, precision, recall, and F1 scores.
Results:
In this study, 538 photos of abdominal ostomies were sent in by 191 patients (median age 52 years; interquartile range, 40-63; 60.4% women). Expert consensus triaged 236 photos (43.9%) as having an obstructed view, 30 photos (5.6%) as healthy, 177 photos (32.9%) for conservative management, and 95 photos (17.7%) for in-person review. All models achieved an area under the receiver operating characteristic curve greater than 0.97 for identifying stomas. The contrastive language-image pretraining vision transformer model performed best at triage (macro-area under the receiver operating characteristic curve 0.94 ± 0.06; F1 score 0.77 ± 0.10). End-to-end detection and triage using contrastive language-image pretraining vision transformer achieved an area under the receiver operating characteristic curve of 0.99 ± 0.01, precision of 0.87 ± 0.08, recall of 0.93 ± 0.04, and F1 score of 0.89 ± 0.07. Attention maps showed that the models focused on stomas to determine classification.
Limitations:
Low sample size and lack of prospective validation.
Conclusions:
A vision-language model pipeline accurately detected and triaged patient-submitted ostomy photos. Prospective evaluation is now needed to support integration into multidisciplinary digital workflows. See Video Abstract .
Enfoque De Visin Artificial Para La Clasificacin De Fotografas De Ostomas Intestinales Enviadas Por Los Pacientes:
ANTECEDENTES:Los cirujanos y el personal de enfermería especializado en ostomías reciben un gran volumen de fotografías de estomas de los pacientes.OBJETIVO:Este estudio tuvo como objetivo desarrollar y validar un sistema automatizado y escalable de inteligencia artificial para la identificación y clasificación de fotografías de ostomías intestinales enviadas por los pacientes.DISEÑO:Estudio retrospectivo de derivación y validación con la participación de las partes interesadas para guiar el diseño del modelo. Los equipos clínicos revisaron las fotografías por duplicado y las clasificaron en: (a) visión obstruida, (b) estoma sano, (c) apto para tratamiento conservador o (d) que requiere revisión presencial. Se optimizaron redes neuronales preentrenadas, incluyendo MobileNetV4, ResNet50, ViT y CLIP-ViT, mediante validación cruzada de 5 pliegues.ÁMBITO:Nueve hospitales de la Clínica Mayo (2019-2022).PACIENTES:Pacientes adultos (≥18 años) sometidos a cirugía que enviaron imágenes dentro de los 30 días posteriores a la intervención.PRINCIPALES MEDIDAS DE RESULTADO:El rendimiento del modelo se evaluó mediante el área bajo la curva ROC, la precisión, la exhaustividad y la puntuación F1.RESULTADOS:En este estudio, 538 fotografías de ostomías abdominales fueron enviadas por 191 pacientes (edad media 52 años, rango intercuartílico 40-63; 60,4 % mujeres). El consenso de expertos clasificó 236 (43,9 %) fotografías como con visión obstruida, 30 (5,6 %) como sanas, 177 (32,9 %) para tratamiento conservador y 95 (17,7 %) para revisión presencial. Todos los modelos lograron un área bajo la curva ROC >0,97 para la identificación de estomas. El modelo CLIP-ViT tuvo el mejor rendimiento en la clasificación (macro: área bajo la curva ROC 0,94 ± 0,06; F1 0,77 ± 0,10). La detección y clasificación de extremo a extremo mediante CLIP-ViT alcanzó un área bajo la curva ROC de 0,99 ± 0,01, una precisión de 0,87 ± 0,08, una exhaustividad de 0,93 ± 0,04 y una puntuación F1 de 0,89 ± 0,07. Los mapas de atención mostraron que los modelos se centraron en los estomas para determinar la clasificación.LIMITACIONES:Tamaño de muestra reducido y falta de validación prospectiva.CONCLUSIONES:Un modelo de lenguaje visual detectó y clasificó con precisión las fotos de ostomías enviadas por los pacientes. Ahora se requiere una evaluación prospectiva para respaldar su integración en flujos de trabajo digitales multidisciplinarios. (AI-generated translation ).
Related Concept Videos
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and solid...
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
