Related Experiment Video
Updated: Apr 30, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Meta-analysis of the accuracy of tools used for binary classification when the primary studies employ different
Juan Botella1, Huiling Huang1, Manuel Suero1
1Department of Social Psychology and Methodology, Universidad Autónoma de Madrid.
Abstract:
The quality of tools used in binary classification is evaluated by studies that assess the accuracy of the classification. The empirical evidence is summarized in 2 × 2 contingency tables. These provide the joint frequencies between the true status of a sample and the classification made by the test. The accuracy of the test is better estimated in a meta-analysis that synthesizes the results of a set of primary studies. The true status is determined by a reference that ideally is a gold standard, which means that it is error free. However, in psychology, it is rare that all the primary studies have employed the same reference, and often they have used an imperfect reference with suboptimal accuracy instead of an actual gold standard. An imperfect reference biases both the estimates of the accuracy of the test and the empirical prevalence of the target status in the primary studies. We discuss several strategies for meta-analysis when different references are employed. Special attention is paid to the simplest case, where the meta-analyst has 1 group of primary studies using a reference that can be considered a gold standard and a 2nd group of primary studies using an imperfect reference. A procedure is recommended in which the frequencies from the primary studies with the imperfect reference are corrected prior to the meta-analysis itself. Then, a hierarchical meta-analytic model is fitted. An example with actual data from SCOFF (Sick-Control-One-Fat-Food; Hill, Reid, Morgan, & Lacey, 2010; Morgan, Reid, & Lacey, 1999) a simple but efficient test for detecting eating disorders, is described.
Related Concept Videos
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Receiver Operating Characteristic Plot
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Comparing the Survival Analysis of Two or More Groups
