Related Experiment Video
Updated: Oct 25, 2025

Evaluating Usability Aspects of a Mixed Reality Solution for Immersive Analytics in Industry 4.0 Scenarios
Published on: October 6, 2020
The Utility of Peers and Trained Raters in Technical Skill-based Assessments a Generalizability Theory Study
Tiffany N Anderson1, James N Lau1, Robert Shi1
1Department of Surgery, Stanford University School of Medicine, Stanford, California.
Faculty surgeons can be assessed by near-peer resident raters (RR) and trained raters (TR) with minimal differences. Training is crucial for optimal procedural competence evaluation, especially for stapled anastomoses.
Area of Science:
- Surgical Education
- Medical Assessment
- Procedural Competence
Background:
- Faculty surgeon shortages challenge resident procedural competence evaluation.
- Validated assessments are the gold standard but difficult to provide.
- Near-peer resident raters (RR), faculty raters (FR), and trained raters (TR) were investigated as alternatives.
Purpose of the Study:
- To determine if procedural assessments by RR, FR, and TR show minimal differences.
- To evaluate the reliability and generalizability of assessments using different rater types.
- To project the number of raters needed for reliable procedural competence evaluation.
Main Methods:
- Deidentified videos of hand-sewn (HA) and stapled (SA) anastomoses were assessed by blinded RR, FR, and TR.
- Intra-class correlation (ICC) and generalizability (g-coefficient) studies were performed.
- Decision study (D-study) projected rater numbers for a g-coefficient > 0.70.
Main Results:
- High ICC values were observed for HA across all rater types (0.84-0.89).
- SA showed lower ICC values for RR (0.77) and FR (0.79) compared to TR (0.86).
- G-coefficients indicated sufficient reliability for HA with 2 raters, but SA required >11 FR for similar reliability.
Conclusions:
- Untrained faculty raters demonstrate high assessment agreement.
- Near-peer resident raters are a viable option for formative surgical skills assessment.
- Training on assessment tools is essential for optimal evaluation of procedural competence, regardless of rater type.
Related Concept Videos
Reliability and Validity
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Theory of Attribution II: Kelley's Covariation Theory
Fundamental Attribution Error
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...

