Bias in Epidemiological Studies
Bias
Reliability and Validity
Testing a Claim about Standard Deviation
Systematic Error: Methodological and Sampling Errors
Statistical Analysis: Overview
You might also read
Articles linked to this work by shared authors, journal, and citation graph.
Updated: Mar 30, 2026

Accuracy in Dental Medicine, A New Way to Measure Trueness and Precision
Published on: April 29, 2014
Ian Schiller1, Maarten van Smeden2, Alula Hadgu3
1Division of Clinical Epidemiology, McGill University Health Centre, Montreal, Canada.
This study examines how combining multiple imperfect diagnostic tests into a single reference standard affects diagnostic accuracy estimates. The researchers found that while adding more tests increases sensitivity, it reduces specificity unless all tests are perfectly specific. This can lead to biased accuracy estimates for new diagnostic tests. The study uses Chlamydia trachomatis testing as an example, where no gold-standard test exists. The findings suggest that commonly used composite reference standards may not reliably improve diagnostic test evaluation. The researchers recommend developing more realistic statistical models to address these limitations.
Area of Science:
Background:
Diagnostic accuracy studies often face limitations due to the absence of a definitive reference standard. In such cases, composite reference standards (CRSs) are proposed as alternatives. These CRSs combine results from multiple imperfect diagnostic tests to classify disease status. While CRSs are intended to improve diagnostic accuracy, they may introduce bias. Prior research has shown that combining multiple tests can enhance sensitivity but may reduce specificity. However, the extent to which CRSs affect accuracy estimates of new tests remains unclear. This gap motivated a detailed analysis of how CRSs influence sensitivity, specificity, and prevalence estimates. The lack of a gold-standard test for asymptomatic diseases like Chlamydia trachomatis highlights the need for a clearer understanding of CRS limitations. Researchers have not yet resolved how conditional dependencies between tests affect accuracy estimates. This uncertainty drives the need for a more rigorous evaluation of CRS-based methods. Understanding these biases is essential for improving diagnostic test evaluation protocols.
Purpose Of The Study:
This study aimed to evaluate the impact of composite reference standards (CRSs) on diagnostic accuracy estimates. The primary goal was to derive algebraic expressions for sensitivity and specificity of CRSs and index tests. The study focused on a CRS that classifies subjects as disease positive if at least one component test is positive. Researchers sought to quantify how CRSs influence accuracy estimates of new tests. The motivation came from the lack of a gold-standard test for asymptomatic diseases like Chlamydia trachomatis. The study sought to clarify how CRSs affect sensitivity, specificity, and prevalence estimates. The researchers aimed to identify conditions under which CRSs introduce bias. This analysis provides a framework for understanding diagnostic test evaluation limitations.
Main Methods:
The study used algebraic derivations to calculate sensitivity and specificity of composite reference standards (CRSs). The CRS classified subjects as disease positive if at least one component test was positive. Researchers derived expressions for sensitivity and specificity of both the CRS and the index test. They also calculated CRS-based prevalence estimates. The analysis focused on a CRS composed of multiple imperfect tests. The study incorporated conditional dependence between the CRS and the index test. Researchers used Chlamydia trachomatis testing as a motivating example. The mathematical framework allowed for evaluating how CRSs influence diagnostic accuracy estimates.
Main Results:
The study found that sensitivity of a composite reference standard (CRS) increases with more component tests. However, specificity decreases unless all tests have perfect specificity. The CRS-based prevalence estimates also changed with increasing number of tests. The accuracy estimates of the index test became significantly biased under these conditions. Conditional dependence between the CRS and index test led to overestimation of accuracy. The bias varied with disease prevalence and CRS accuracy. The study showed that CRSs may not improve over single imperfect tests unless specific conditions are met. These findings highlight limitations in commonly used CRS approaches.
Conclusions:
The study demonstrated that composite reference standards (CRSs) can introduce bias in diagnostic accuracy estimates. The CRS sensitivity increases with more component tests but at the expense of specificity. The index test accuracy estimates become biased unless all component tests have perfect specificity. Conditional dependence between the CRS and index test further exacerbates this bias. The study showed that CRSs may not improve over single imperfect tests in most cases. The findings suggest that current CRS approaches may not be reliable for diagnostic test evaluation. The study emphasizes the need for alternative statistical models in the absence of gold-standard tests. Researchers propose that more realistic models should be developed for diagnostic accuracy studies.
CRSs may increase sensitivity but reduce specificity unless all component tests have perfect specificity.
More component tests increase CRS sensitivity but decrease specificity unless all tests are perfectly specific.
Conditional dependence can lead to overestimation of index test accuracy when using a CRS.
Bias in accuracy estimates depends on both disease prevalence and the accuracy of the CRS.
Only if all component tests have perfect specificity and the CRS is conditionally independent of the index test.
The authors propose developing more realistic statistical models instead of relying on CRSs.