Related Experiment Video
Updated: Feb 7, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Interobserver agreement between an artificial intelligence algorithm and colon capsule endoscopy readers on
Benedicte Schelde-Olesen1,2, Jürgen Herp3,4, Jan-Matthias Braun3,4
1Department of Clinical Research, University of Southern Denmark, Odense, Denmark.
Background And Aims:
Colon capsule endoscopy (CCE) faces substantial challenges, one of which is achieving adequate colon cleansing. Furthermore, the interobserver agreement on bowel-cleansing quality varies. To address this issue, we developed an artificial intelligence algorithm (AIA) to evaluate bowel-cleansing quality. The aim of this study was to estimate the interobserver agreement on bowel cleansing between a group of experienced CCE readers and an AIA and to examine whether percentiles of the overall bowel-cleansing quality are a suitable way of reporting the results generated by the AIA.
Methods:
Bowel-cleansing quality in 842 CCE investigations was scored on both a 2- and 4-point grading scale for the entire colon and by segment by experienced CCE readers and the AIA. For the algorithm, a score was given based on the mean score, median, upper and lower quartiles, and second and 98th percentiles. The level of agreement was evaluated using Cohen's κ.
Results:
The interobserver agreement between the CCE readers and AIA on bowel-cleansing quality was minimal to none for the overall bowel evaluation, by segment, and on the 2- and 4- point grading scale regardless of the threshold for the AIA score.
Conclusions:
We found minimal agreement on evaluation of bowel-cleansing quality in CCE between CCE readers and the AIA. Mean or percentiles of the AIA grading did not seem suitable for AI-generated bowel-cleansing evaluation.
Related Concept Videos
Endoscopic Procedures III: Video Capsule Endoscopy
The Colonization of Land
Intelligence
Trial and Error and Algorithm
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Multiple Intelligences Theory

