Related Experiment Video
Updated: Oct 25, 2025

Evaluating Usability Aspects of a Mixed Reality Solution for Immersive Analytics in Industry 4.0 Scenarios
Published on: October 6, 2020
The Utility of Peers and Trained Raters in Technical Skill-based Assessments a Generalizability Theory Study
Tiffany N Anderson1, James N Lau1, Robert Shi1
1Department of Surgery, Stanford University School of Medicine, Stanford, California.
Objective:
The gold standard for evaluation of resident procedural competence is that of validated assessments from faculty surgeons. A provision of adequate trainee assessments is challenged by a shortage of faculty due to increased clinical and administrative responsibilities. We hypothesized that with a well constructed assessment instrument and training, there would be minimal differences in procedural assessments made by near-peer resident raters (RR), faculty raters (FR), and trained raters (TR).
Design:
Deidentified videos of residents performing hand-sewn (HA) and stapled (SA) anastomoses were distributed to blinded reviewers of 3 types. Intra-class correlation (ICC) of RR, FR and TR assessments was determined for each procedure. A fully-crossed design was used to examine the internal structure validity in a generalizability study. A Decision study was performed to make projections on the number of raters needed for a g-coefficient > 0.70.
Setting:
This study was conducted within a private academic institution, using the creation of intestinal anastomoses as the procedural model.
Participants:
Raters consisted of residents who were untrained to the assessment (UTA) tool, UTA faculty surgeons, and individuals with training.
Results:
Twenty nine videos were reviewed (15 HA and 14 SA) by a total of 9 video reviewers (4 RR, 2 FR, and 3 TR). HA ICC values were 0.84 (Confidence Interval [CI]:0.81-0.87) for RR, 0.89 (CI:0.86-0.92) for FR, and 0.88 (CI:0.86-0.90) for TR. SA ICC values were 0.77 (CI:0.72-0.80) for RR, 0.79 (CI:0.75-0.83) for FR, and 0.86 (CI:0.83-0.88) for TR. The g-coefficient was RR = 0.72, FR = 0.85, and TR = 0.77 for HA; and RR = 0.33, FR = 0.38, and TR = 0.4 for SA. The D-study indicated that at least 2 raters of any type were needed for HA and > 11 FR for SA.
Conclusions:
Faculty without training have high assessment agreement. Peers for surgical skills assessment is an option for formative evaluation without training. Training to assessment tools should be performed for any assessment, formative or summative, for the optimal evaluation of procedural competence.
Related Concept Videos
Reliability and Validity
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Theory of Attribution II: Kelley's Covariation Theory
Fundamental Attribution Error
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...

