Related Experiment Video
Updated: May 3, 2026

Meta-analysis of Voxel-Based Neuroimaging Studies using Seed-based d Mapping with Permutation of Subject Images SDM-PSI
Published on: November 27, 2019
A simulation study into the performance of "optimal" diagnostic thresholds in the population:"Large" effect sizes are
Gerrit Hirschfeld1, Pedro Emmanuel Alvarenga Americano do Brasil2
1German Pediatric Pain Center, Children's Hospital, Dr.-Friedrich-Steiner Str. 5, 45711 Datteln, Germany; Children's Pain Therapy and Paediatric Palliative Care, Witten/Herdecke University, 45711 Datteln, Germany.
Objectives:
Many diagnostic studies are aimed at defining "optimal" thresholds. Here, we evaluate the performance of empirically defined optimal thresholds (1) in the sample in which they were defined and (2) in the population from which the sample was drawn.
Study Design And Setting:
We simulated test results for 120,000 samples varying the number of people without a disease (n between 20 and 500), number of people with a disease (m between 20 and 500), the magnitude of the difference between group means [effect size (ES) between 0.5 and 4], and distributions (normal and log-normal). The thresholds associated with the maximal Youden index were defined as optimal. Performance was defined as the percentage of correct classifications in the sample and when applied to the whole population.
Results:
At the population level, the thresholds defined for the four ESs (0.5, 0.8, 2, and 4) yielded a median of 59%, 65%, 83%, and 97% correct classifications, respectively. At the sample level, the samples with similar characteristics yielded widely varying estimates of the performance that were systematically higher than at the population level.
Conclusion:
Researchers need to be careful defining cut points for mean differences that are traditionally considered "large" (ES = 0.8). The diagnostic utility of optimal thresholds needs to be assessed in prospective studies.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Receiver Operating Characteristic Plot
Regression Toward the Mean
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Errors In Hypothesis Tests

