Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

1.6K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
1.6K
Bias01:22

Bias

8.0K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
8.0K
Reliability and Validity01:29

Reliability and Validity

14.4K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.4K
Testing a Claim about Standard Deviation01:19

Testing a Claim about Standard Deviation

3.2K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
3.2K
Systematic Error: Methodological and Sampling Errors01:15

Systematic Error: Methodological and Sampling Errors

11.4K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
11.4K
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

17.2K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
17.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Intranasal monoclonal antibodies do not prevent respiratory infection in a randomized, controlled experimental infection trial.

npj drug discovery·2026
Same author

Estimating diagnostic accuracy under uncertainty about disease status: a sepsis case study.

Journal of clinical epidemiology·2026
Same author

Circulating Tumour DNA and Extracellular Vesicle-Associated DNA as Biomarkers for Cancer Detection: A Systematic Review and Meta-Analysis.

Journal of extracellular biology·2026
Same author

QUADAS-3 Explanation and Elaboration: Guidance for Quality Assessment of Diagnostic Test Accuracy Studies.

Annals of internal medicine·2026
Same author

QUADAS-3: A Revised Tool for the Quality Assessment of Diagnostic Test Accuracy Studies.

Annals of internal medicine·2026
Same author

Real-world Inter-rater Agreement of PI-QUAL Version 2 for Prostate Magnetic Resonance Imaging Quality Assessment and Its Association with Diagnostic Accuracy.

European urology open science·2026

Related Experiment Video

Updated: Mar 30, 2026

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
07:57

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision

Published on: April 29, 2014

14.1K

Bias due to composite reference standards in diagnostic accuracy studies.

Ian Schiller1, Maarten van Smeden2, Alula Hadgu3

  • 1Division of Clinical Epidemiology, McGill University Health Centre, Montreal, Canada.

Statistics in Medicine
|November 12, 2015
PubMed
Summary

This study examines how combining multiple imperfect diagnostic tests into a single reference standard affects diagnostic accuracy estimates. The researchers found that while adding more tests increases sensitivity, it reduces specificity unless all tests are perfectly specific. This can lead to biased accuracy estimates for new diagnostic tests. The study uses Chlamydia trachomatis testing as an example, where no gold-standard test exists. The findings suggest that commonly used composite reference standards may not reliably improve diagnostic test evaluation. The researchers recommend developing more realistic statistical models to address these limitations.

Keywords:
compositeconditional dependenceimperfect referencesensitivityspecificitycomposite reference standarddiagnostic test evaluationChlamydia trachomatis testingdiagnostic accuracy bias

Frequently Asked Questions

More Related Videos

Assessment of Child Anthropometry in a Large Epidemiologic Study
09:36

Assessment of Child Anthropometry in a Large Epidemiologic Study

Published on: February 2, 2017

28.0K
Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia
06:19

Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia

Published on: September 27, 2024

631

Related Experiment Videos

Last Updated: Mar 30, 2026

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
07:57

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision

Published on: April 29, 2014

14.1K
Assessment of Child Anthropometry in a Large Epidemiologic Study
09:36

Assessment of Child Anthropometry in a Large Epidemiologic Study

Published on: February 2, 2017

28.0K
Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia
06:19

Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia

Published on: September 27, 2024

631

Area of Science:

  • Medical diagnostics
  • Epidemiological methods
  • Infectious disease testing

Background:

Diagnostic accuracy studies often face limitations due to the absence of a definitive reference standard. In such cases, composite reference standards (CRSs) are proposed as alternatives. These CRSs combine results from multiple imperfect diagnostic tests to classify disease status. While CRSs are intended to improve diagnostic accuracy, they may introduce bias. Prior research has shown that combining multiple tests can enhance sensitivity but may reduce specificity. However, the extent to which CRSs affect accuracy estimates of new tests remains unclear. This gap motivated a detailed analysis of how CRSs influence sensitivity, specificity, and prevalence estimates. The lack of a gold-standard test for asymptomatic diseases like Chlamydia trachomatis highlights the need for a clearer understanding of CRS limitations. Researchers have not yet resolved how conditional dependencies between tests affect accuracy estimates. This uncertainty drives the need for a more rigorous evaluation of CRS-based methods. Understanding these biases is essential for improving diagnostic test evaluation protocols.

Purpose Of The Study:

This study aimed to evaluate the impact of composite reference standards (CRSs) on diagnostic accuracy estimates. The primary goal was to derive algebraic expressions for sensitivity and specificity of CRSs and index tests. The study focused on a CRS that classifies subjects as disease positive if at least one component test is positive. Researchers sought to quantify how CRSs influence accuracy estimates of new tests. The motivation came from the lack of a gold-standard test for asymptomatic diseases like Chlamydia trachomatis. The study sought to clarify how CRSs affect sensitivity, specificity, and prevalence estimates. The researchers aimed to identify conditions under which CRSs introduce bias. This analysis provides a framework for understanding diagnostic test evaluation limitations.

Main Methods:

The study used algebraic derivations to calculate sensitivity and specificity of composite reference standards (CRSs). The CRS classified subjects as disease positive if at least one component test was positive. Researchers derived expressions for sensitivity and specificity of both the CRS and the index test. They also calculated CRS-based prevalence estimates. The analysis focused on a CRS composed of multiple imperfect tests. The study incorporated conditional dependence between the CRS and the index test. Researchers used Chlamydia trachomatis testing as a motivating example. The mathematical framework allowed for evaluating how CRSs influence diagnostic accuracy estimates.

Main Results:

The study found that sensitivity of a composite reference standard (CRS) increases with more component tests. However, specificity decreases unless all tests have perfect specificity. The CRS-based prevalence estimates also changed with increasing number of tests. The accuracy estimates of the index test became significantly biased under these conditions. Conditional dependence between the CRS and index test led to overestimation of accuracy. The bias varied with disease prevalence and CRS accuracy. The study showed that CRSs may not improve over single imperfect tests unless specific conditions are met. These findings highlight limitations in commonly used CRS approaches.

Conclusions:

The study demonstrated that composite reference standards (CRSs) can introduce bias in diagnostic accuracy estimates. The CRS sensitivity increases with more component tests but at the expense of specificity. The index test accuracy estimates become biased unless all component tests have perfect specificity. Conditional dependence between the CRS and index test further exacerbates this bias. The study showed that CRSs may not improve over single imperfect tests in most cases. The findings suggest that current CRS approaches may not be reliable for diagnostic test evaluation. The study emphasizes the need for alternative statistical models in the absence of gold-standard tests. Researchers propose that more realistic models should be developed for diagnostic accuracy studies.

CRSs may increase sensitivity but reduce specificity unless all component tests have perfect specificity.

More component tests increase CRS sensitivity but decrease specificity unless all tests are perfectly specific.

Conditional dependence can lead to overestimation of index test accuracy when using a CRS.

Bias in accuracy estimates depends on both disease prevalence and the accuracy of the CRS.

Only if all component tests have perfect specificity and the CRS is conditionally independent of the index test.

The authors propose developing more realistic statistical models instead of relying on CRSs.