Related Experiment Video
Updated: Jan 17, 2026

Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Evaluating the evaluators: does C-SATS measure up?
Robert B Laverty1, Charles H Chesnut2, Joseph R Karam3
1Department of Surgery, Brooke Army Medical Center, San Antonio, TX, USA. rblaverty@gmail.com.
Introduction:
Robotic-assisted surgery has increased in prevalence, particularly in general surgery. The number of cases required to achieve adequate proficiency in robotic surgery, however, and the training metrics that correlate best with proficiency remain unclear. We sought to better define proficiency-based benchmarks in robotic-assisted cholecystectomies (RAC) and inguinal hernia repairs (RIHR) using a commercial crowd source based on competency platform.
Methods:
Multi-institutional cohort study in which 48 surgeons (senior residents, fellows, and practicing physicians) submitted representative videos of themselves performing a RAC and/or RIHR. Subjects subsequently underwent blinded case video reviews using the C-SATS platform, which utilizes the Global Evaluative Assessment of Robotic Skills (GEARS) rubric. Participating surgeons self-reported surgical case volume. Primary outcome was correlation of GEARS scores with historic procedure case volume. Secondary outcomes included construct validity of GEARS scores as an operative proficiency metric.
Results:
Total GEARS scores and historical case volume showed positive correlation for both RAC (r = 0.65, p < 0.0001) and RIHR (r = 0.54, p = 0.001) among all performers. On subgroup analysis, no correlation was seen for resident/fellow physicians (r = 0.39, p = 0.11 for RAC; r = 0.22, p = 0.49 for RIHR) or those with < 50 historic case volume (r = 0.14, p = 0.55 for RAC; r = 0.21, p = 0.54 for RIHR). No difference in total GEARS scores was seen between resident/fellow and practicing physicians for either RAC (20.21 v 20.25, p = 0.82) or RIHR (20.45 v 20.46, p = 0.95), nor in those with < 50 or ≥ 50 historic case volume in RAC (20.16 v 20.33, p = 0.33) and RIHR (20.35 v 20.49, p = 0.48). GEARS scores by domain (bimanual dexterity, depth perception, efficiency, force sensitivity, and robotic control) and surgical step (exposure of triangle of calot, clipping and division of cystic artery/duct, and dissection of gallbladder; mobilizing peritoneal flap, hernia sac dissection, and mesh placement) were similar across both groups (p > 0.05).
Conclusion:
C-SATS-derived GEARS scores correlated to overall surgeon historical case volume for RA cholecystectomy and IHR, but not among novice performers. This methodology was unable to differentiate between novice and expert performers for these procedures. There remains a need for high-fidelity and discerning robotic skills evaluation platforms for trainees and novice surgeons.
Related Concept Videos
Reliability and Validity
SI Units: 2019 Redefinition
A standard set of units has been defined...
Statistical Analysis System (SAS)
Applications: SAS finds applications in numerous fields, including healthcare for clinical trial analysis, finance for risk assessment, marketing for customer data analysis, and...
Statgraphics
Introduction to Statistical Process Control
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...

