Related Experiment Video
Updated: Aug 8, 2026

Measurement of Specific Mycobacterial Mistranslation Rates with Gain-of-function Reporter Systems
Published on: April 26, 2019
Evaluating the applicability of replication success metrics in animal-to-human translation: A simulation study
Carolyne Jie Huang1,2, Samuel Pawel3,4, Kimberley Elaine Wever5
1Master Program in Biostatistics, Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland.
Abstract:
Translation failure, in which promising animal study results cannot be reproduced in human trials, is a challenge in biomedical research. Metrics for replication success are widely used to evaluate reproducibility, that is the extent to which the results of a study agree with those of replication studies. The relevance of these metrics in assessing animal-to-human translation success (or failure) is unclear. We conducted a simulation study to examine whether these metrics can quantify translation success, and how their performance varies under different conditions. Using parameters from a meta-analysis on prenatal amino acid supplementation and maternal blood pressure, we simulated animal and human studies under 648 scenarios, varying effect sizes, heterogeneity, animal sample sizes, and number of pooled animal studies. Nine metrics were assessed, namely the two-trials rule, meta-analysis, replication Bayes factor, unweighted and weighted Edgington's methods, golden skeptical p-value, and three versions of controlled skeptical p-value. Most metrics, except meta-analysis and replication Bayes factor, controlled false positive rates under no heterogeneity, but became liberal as heterogeneity increased, particularly between human studies. Translation power (i.e. the probability of true positive translation success) was constrained by the weaker evidence of the two findings; for example, small sample size in the animal studies resulted in lower translation power. The metric based on meta-analysis frequently indicated success when either of the species found strong evidence, while skeptical p-values were more conservative. The skeptical p-value that controls overall type-one error and the weighted version of Edgington's method performed relatively consistently across scenarios. However, no metric was uniformly optimal. Metrics developed for replication studies can inform assessments of translation, but their utility depends on the underlying evidence and assumptions. Using multiple metrics in combination, with attention to their strengths and limitations, is recommended for evaluating the translation of animal findings to human outcomes.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Replication in Eukaryotes
Leaky Scanning
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Replication in Prokaryotes

