Related Experiment Video
Updated: Mar 10, 2026

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
Misusing I2 for inconsistency, overlooking OIS for imprecision, and ignoring the continuum of certainty ratings:
Prashanti Eachempati1, Gordon Guyatt2
1MAGIC Evidence Ecosystem Foundation, Oslo, Norway; Peninsula Dental School, University of Plymouth, United Kingdom; Faculty of Dentistry, Manipal University College Malaysia, Melaka, Malaysia.
Background:
The GRADE framework guides ratings of certainty in evidence that includes the possibility of rating down certainty in one or more of five domains. GRADE users face challenges in making certainty of evidence judgments, and these challenges can result in inappropriate ratings. Potential problems include over-reliance on the I² statistic, neglect of sample size adequacy, and binary decision-making.
Objectives:
To demonstrate applying key principles of visual criteria for inconsistency judgments, sample size considerations for imprecision judgments, and domain and overall certainty judgments on a continuum.
Methods:
We use examples from four meta-analyses, one evaluating nasal continuous positive airway pressure (NCPAP) in preterm infants, two comparing anti-malarial regimens, and another comparing tooth brushing versus no tooth brushing for ventilator-associated pneumonia. These examples illustrate the application of key principles, including visual criteria, sample size considerations, and certainty judgments along a continuum.
Results:
In two examples, I² was high, but visual inspection showed consistent point estimates, overlapping confidence intervals, and estimates all on the same side of the threshold. Therefore, rating down for inconsistency was not justified. In two other examples, we calculated the Optimal Information Size (OIS) using a 25% relative risk reduction. In both, the sample size fell short, warranting a rating down for imprecision. In one antimalarial efficacy example, the sample size failed to meet the OIS based on 25% RRR but would have been sufficient with a 30% RRR. This placed the concern at the lower end of serious. These domain-level judgments, made along a continuum, contributed to overall certainty ratings that lay at the upper end of their respective categories.
Conclusions:
Applying GRADE principles through visual inspection of forest plots in evaluation of consistency, OIS-based evaluation of precision, and continuum-based domain judgments supports more transparent, and decision-relevant certainty ratings. Clearly describing the degree of concern within each domain enhances the clarity and utility of evidence summaries for decision-makers.
Related Concept Videos
Common Leveling Mistakes and Errors
Uncertainty in Measurement: Reading Instruments
Inductively Coupled Plasma-Mass Spectrometry (ICP-MS): Interferences
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Accuracy and Precision
Errors In Hypothesis Tests

