Related Experiment Videos
Bayesian and fuzzy hybrid standard setting for pass-fail decisions in clinical examinations
Yogeesh N1,2,3, Mohammed Almakki4, Asokan Vasudevan5,6,7
1School of Engineering, Architecture and Interior Design, Amity University Dubai, Dubai International Academic City, Dubai, 345019, United Arab Emirates. yogeesh.r@gmail.com.
Background:
Pass-fail decisions in clinical examinations must be defensible, yet traditional standard-setting approaches often report a single cut score without explicitly quantifying uncertainty-an issue amplified in small cohorts and mixed-format assessments (e.g., OSCE plus written). We propose a Bayesian fuzzy hybrid standard-setting framework that (i) treats station cut scores as posterior distributions and (ii) models the inherently linguistic "borderline" construct using fuzzy membership, yielding both a central standard and a principled borderline review band.
Methods:
A mixed-format assessment model was specified with total score [Formula: see text], where [Formula: see text] is an equally weighted mean of [Formula: see text] OSCE stations with global ratings [Formula: see text]. Station cut scores were estimated using (a) Borderline Regression Method (BRM) for comparison and (b) Bayesian regression [Formula: see text], giving posterior [Formula: see text] at borderline [Formula: see text]. Borderline semantics were represented via a trapezoidal fuzzy set [Formula: see text] with centroid [Formula: see text], producing fuzzy-adjusted cuts [Formula: see text]. A hybrid OSCE standard [Formula: see text] was mapped to the mixed-format cut [Formula: see text]. A decision band [Formula: see text] combined Bayesian credible uncertainty and fuzzy [Formula: see text]-cut uncertainty to classify candidates as Pass, Fail, or Borderline review. The full workflow was demonstrated using a collected sample dataset (n=34, k=6) to illustrate reporting and reproducibility.
Results:
Station-level BRM cut scores ranged from 59.5 to 62.4, while Bayesian station cut scores produced interpretable 95% credible intervals around similar means. The Bayesian OSCE cut score mean was approximately 61.2 with a narrow posterior interval, and the hybrid mixed format cut score was approximately 60.8 (0-100 scale). The uncertainty-aware decision band produced 22 Pass (64.7%), 10 Fail (29.4%), and 2 Borderline review (5.9%) classifications, explicitly isolating boundary cases rather than forcing deterministic decisions. Bootstrap resampling indicated that the hybrid central standard was stable and comparable to BRM, while adding transparency via a bounded review zone.
Conclusions:
Here the Bayesian-fuzzy hybrid framework retained the interpretability of borderline regression but provided explicit uncertainty quantification and a structured bandfor borderline review. The results ought to be understood as methodological and illustrative but not confirmation of universal superiority; external validation in larger multi-centre clinical examination cohorts is necessary prior to routine operational implementation.
Related Concept Videos
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Expected Frequencies in Goodness-of-Fit Tests
Hypothesis: Accept or Fail to Reject?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null hypothesis and 'fail to...
Errors In Hypothesis Tests