Back

Reply to: Caution Regarding the Specificities of Pan-Cancer Microbial Structure

Sepich-Poore, G. D.; Kopylova, E.; Zhu, Q.; Carpenter, C.; Fraraccio, S.; Wandro, S.; Kosciolek, T.; Janssen, S.; Metcalf, J.; Song, S. J.; Kanbar, J.; Miller-Montgomery, S.; Heaton, R.; Mckay, R.; Patel, S. P.; Swafford, A. D.; Knight, R.

2023-02-13 bioinformatics
10.1101/2023.02.10.528049 bioRxiv
Show abstract

The cancer microbiome field tremendously accelerated following the release of our manuscript nearly three years ago1, including direct validation of our cancer type-specific conclusions in independent, international cohorts2,3 and the tumor microbiomes adoption into the hallmarks of cancer4. Disentangling contamination signals from biological signals is an important consideration for this research field. Therefore, despite numerous, high-impact, peer-reviewed research papers that either validated our conclusions or extended them using data we released2,5-13, we carefully considered criticism raised by Gihawi et al. about potential mishandling of contaminants, batch effects, and machine learning approaches--all of which were central topics in our manuscript. Nonetheless, a close examination of each concern alongside the original manuscript and re-analyses of our published data strongly demonstrates the robustness of the original findings. To remove all doubt, however, we have reproduced all key conclusions from the original manuscript using only overlapping bacterial genera identified in a highly decontaminated, multi-cancer, international cohort (Weizmann Institute of Science, WIS)2, with or without batch correction, and with multiclass machine learning analyses to mitigate class imbalances. Our published pan-cancer mycobiome manuscript3 also affirms these findings using updated, state-of-the-art methods. We also note that every analysis shown here was possible using public data and code that we had already provided.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.