Evaluating the applicability of replication success metrics in animal-to-human translation: A simulation study
Huang, C. J.; Pawel, S.; Wever, K. E.; Ineichen, B. V.; Heyard, R.
Show abstract
Translation failure, in which promising animal study results can not be reproduced in human trials, is a challenge in biomedical research. Metrics for replication success are widely used to evaluate reproducibility, i.e., the extent to which the results of a study agree with those of replication studies. The relevance of these metrics in assessing animal-to-human translation success (or faillure) is unclear. We conducted a simulation study to examine whether these metrics can quantify translation success and how their performance varies under different conditions. Using parameters from a meta-analysis on prenatal amino acid supplementation and maternal blood pressure, we simulated animal and human studies under 648 scenarios, varying effect sizes, heterogeneity, animal sample sizes and number of pooled animal studies. Nine metrics were assessed, namely the two-trials rule, meta-analysis, replication Bayes factor, unweighted and weighted Edgingtons methods, golden sceptical p-value and three versions of controlled sceptical p-value. Most metrics, except meta-analysis and replication Bayes factor, controlled false positive rates under no heterogeneity, but became liberal as heterogeneity increased, particularly between human studies. Translation power (i.e., the probability of true positive translation success) was constrained by the weaker evidence of the two findings; e.g., small sample size in the animal studies resulted in lower translation power. The metric based on meta-analysis frequently indicated success when either of the species found strong evidence, while sceptical p-values were more conservative. The sceptical p-value that controls overall type-one error and the weighted version of Edgingtons method performed relatively consistently across scenarios. No metric was uniformly optimal. Metrics developed for replication studies can inform assessments of translation, but their utility depends on the underlying evidence and assumptions. Using multiple metrics in combination, with attention to their strengths and limitations, is recommended for evaluating the translation of animal findings to human outcomes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Data-driven Prior Elicitation for Bayes Factors in Cox Regression for Nine Subfields in Biomedicine 95%
- Using numerical modelling and simulation to assess the ethical burden in clinical trials and how it relates to the proportion of responders in a trial sample 94%
- A method of back-calculating the log odds ratio and standard error of the log odds ratio from the reported group-level risk of disease 94%
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 95%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 93%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 93%
Similar papers in this journal
- ZIBGLMM: Zero-Inflated Bivariate Generalized Linear Mixed Model for Meta-Analysis with Double-Zero-Event Studies 95%
- Evaluation of statistical methods used to meta-analyse results from interrupted time series studies: a simulation study 95%
- Synthesizing evidence from the earliest studies to support decision-making: to what extent could the evidence be reliable? 94%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 93%
- Estimating and Testing an Index of Bias Attributable to Composite Outcomes in Comparative Studies 93%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 92%
Similar papers in this journal
- Use of estimands in cluster randomised trials: a review 92%
- Estimating Counterfactual Placebo HIV Incidence in HIV Prevention Trials Without Placebo Arms Based on Markers of HIV Exposure 92%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.