An Empirical Assessment of Inferential Reproducibility of Linear Regression in Health and Biomedical Research Papers
Jones, L.; Barnett, A.; Hartel, G.; Vagenas, D.
Show abstract
Background: In health research, variability in modelling decisions can lead to different conclusions even when the same data are analysed, a challenge known as inferential reproducibility. In linear regression analyses, incorrect handling of key assumptions, such as normality of the residuals and linearity, can undermine reproducibility. This study examines how violations of these assumptions influence inferential conclusions when the same data are reanalysed. Methods: We randomly sampled 95 health-related PLOS ONE papers from 2019 that reported linear regression in their methods. Data were available for 43 papers, and 20 were assessed for computational reproducibility, with three models per paper evaluated. The 14 papers that included a model at least partially computationally reproduced were then examined for inferential reproducibility. To assess the impact of assumption violations, differences in coefficients, 95% confidence intervals, and model fit were compared. Results: Of the fourteen papers assessed, only three were inferentially reproducible. The most frequently violated assumptions were normality and independence, each occurring in eight papers. Violations of independence were particularly consequential and were commonly associated with inferential failure. Although reproduced analyses often retained the same binary statistical significance classification as the original studies, confidence intervals were frequently wider, indicating greater uncertainty and reduced precision. Such uncertainty may affect the interpretation of results and, in turn, influence treatment decisions and clinical practice. Conclusion: Our findings demonstrate that substantial violations of key modelling assumptions often went undetected by authors and peer reviewers and, in many cases, were associated with inferential reproducibility failure. This highlights the need for stronger statistical education and greater transparency in modelling decisions. Rather than applying rigid or misinformed rules, such as incorrectly testing the normality of the outcome variable, researchers should adopt modelling frameworks guided by the research question and the study design. When assumptions are violated, appropriate alternatives, such as robust methods, bootstrapping, generalized linear models, or mixed-effects models, should be considered. Given that assumption violations were common even in relatively simple regression models, early and sustained collaboration with statisticians is critical for supporting robust, defensible, and clinically meaningful conclusions.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 99%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 96%
- A method of back-calculating the log odds ratio and standard error of the log odds ratio from the reported group-level risk of disease 94%
Similar papers in this journal
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 97%
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
- Development, validation, and usage of metrics to evaluate clinical research hypothesis quality 94%
Similar papers in this journal
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 95%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 95%
- Strategies used to manage overlap of primary study data by exercise-related overviews. Protocol for a systematic methodological review 93%
Similar papers in this journal
- Collecting mortality data via mobile phone surveys: a non-inferiority randomized trial in Malawi 92%
- Factors affecting trust in clinical trials conduct: Views of stakeholders from a qualitative study in Ghana 92%
- Self-tests for COVID-19: what is the evidence? A living systematic review and meta-analysis (2020-2023) 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.