Related Experiment Video
Updated: Apr 18, 2026

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
Validity and Reliability of SCOPE (Structured Comprehensive Oral Problem-based Examination) using Generalizability
Mashaal Sabqat1, Noorul Ain2, Rehan Ahmed Khan3
1Mashaal Sabqat, MBBS, MHPE Assistant Professor Medical Education, Islamic International Medical College, Assistant Director Riphah Institute of Assessment, Riphah International University, Islamabad, Pakistan.
Objective:
To validate the SCOPE tool, determine its reliability, and identify sources of score variation using Generalizability theory, including a Decision Study to optimize its structure.
Methodology:
The study was conducted at Islamic International Medical College (IIMC), Rawalpindi from March 2025 to May 2025. SCOPE (Structured Comprehensive Oral Problem-based Examination) marking sheet and CVI (Content Validity Index) form were reviewed by medical educationalists for feedback and validation. Reliability was assessed through a pilot involving 37 final-year medical students. Four trained examiners conducted SCOPE, with each student assessed by one examiner using problems derived from predefined must-know and good-to-know topic lists. Performance was scored using a structured marking sheet for Recall (five marks) and Application (20 marks) categories. Reliability was assessed using Generalizability study (G-Study) with a crossover random-effects design, followed by a Decision Study (D-Study) to estimate projected reliability coefficients across varying numbers of problems and categories.
Results:
Ten medical educationalists provided qualitative feedback and rated the relevance and clarity of SCOPE sheet, demonstrating strong content validity (S-CVI/Ave = 0.92) and clarity (average CCA = 2.85). Reliability analysis showed a G-Coefficient value of 0.793, and Phi-coefficient value of 0.696, indicating good reliability for ranking students, and moderate reliability for absolute decisions. Students contributed the largest variance (48.13%), followed by student × problem × category interaction (19.8%). D-Study revealed that increasing the number of problems and categories substantially improved both relative and absolute reliability.
Conclusion:
SCOPE is a feasible, reliable and innovative tool for comparative assessments in low-resource settings, with reliability further enhanced by increasing problem numbers for each student.
Related Concept Videos
Reliability and Validity
Theory of Attribution II: Kelley's Covariation Theory
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Methods of Documentation III: PIE
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
