Back

Supporting Reanalysis and Reuse of Clinical Trial Data: A Case Study

Burgwinkel, C.; Chiam, H. C.; Tai, K. H.; Wang, J.; Ali, M. H.; Fallah, S. S.; Matbouriahi, M.; Obinwanne, T.; Papapostolou, G.; Riedha, M.; Varvara, G.; Zalai, Y.; Mansmann, U.; Sax, U.; Held, L.

2025-11-07 oncology
10.1101/2025.11.06.25339683 medRxiv
Show abstract

BackgroundReproducing published findings from clinical trials is a critical component of scientific transparency, yet it remains a challenging and under-practiced task. Despite increasing emphasis on reproducibility and data reuse in research policies, few real-world examples exist where independent teams have reproduced complex analyses using clinical trial data. In this case study, the aim was to independently reproduce the key findings of a high-impact clinical trial on rectal cancer treatment using shared trial data. MethodWe organized a multi-team datathon, where each team was provided with the same dataset and supporting material, and was tasked to reproduce the results of the CAO/ARO/AIO-04 trial, with optional additional analysis. We contacted the original investigators for data access and reuse, and consulted them to understand the study, clinically and scientifically. ResultsFive teams used R or Python to reproduce the statistical results, and the corresponding scripts can be found on Gitlab. All teams reproduced the analyses for primary outcome--disease-free survival (DFS). The key findings on DFS were consistently reproduced, reinforcing confidence in the trial main conclusions. Result robustness was investigated using a different analytical software or statistical models. Nevertheless, challenges were encountered when the supplementary materials were not easily identified. Minor reporting issues were noticed in the reproduced paper. ConclusionReproduction of a major oncology clinical trial confirmed the reliability of its main conclusions. Divergences highlighted reporting gaps--such as incomplete protocols and broken links --that future trials should address. This case study demonstrates the value of systematic reproducibility checks for clinical research transparency and challenges in data sharing for reproducibility.

Published in Trials (predicted rank #1) · training set

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.