Related Experiment Video
Updated: Jan 10, 2026

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
CONTEST: A generalization of ONEST to estimate sample size for predictive augmented intelligence method validation
Benjamin K Olson1, Joseph H Rosenthal2, Ryan D Kappedal3
1University of California Santa Cruz, Santa Cruz, CA, USA.
Abstract:
Laboratories must verify and validate assays before reporting results in the clinical record. With the advent of machine learning algorithms, multiclass decision-support tools are coming online but the FDA explicitly does not contemplate multiclass problems in their guidance for test validation. Validation requires, for a laboratory's patient population, evaluation of four performance characteristics to a reference method: accuracy, precision, reportable range, and reference intervals. In the absence of a reference method, proportion of agreement is the appropriate metric (Meier 2007). For subjective tests, the traditional metrics for precision are in the area of interrater reliability, and interrater reliability is well studied in the pathology literature ("Gwet Handbook of Interrater Reliability 4th Ed.pdf," n.d.). Recently, Guo and Han introduced an alternative framing, Observers Needed to Evaluate a Subjective Test (ONEST). This article introduces a treatment effect extension of ONEST, Cases and Observers Needed to Evaluate a Subjective Test (CONTEST) and demonstrates that the agreement and disagreement distributions can be reasonably specified with parametric probability distributions such that the required sample size for a test, at a given level and power, can be calculated. We argue that this would be an appropriate method to develop for validation of tools used to augment a subjective test, given a prior set of cases, observers, and decisions, such as from another archive, cohort, or dataset, particularly in resource-constrained settings.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
One-Way ANOVA: Unequal Sample Sizes
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

