Patterns of 'Analytical Irreproducibility' in Multimodal Diseases
Basson, A. R.; Cominelli, F.; Rodriguez-Palacios, A.
Show abstract
Multimodal diseases are those in which affected individuals can be divided into subtypes (or data modes); for instance, mild vs. severe, based on (unknown) modifiers of disease severity. Studies have shown that despite the inclusion of a large number of subjects, the causal role of the microbiome in human diseases remains uncertain. The role of the microbiome in multimodal diseases has been studied in animals; however, findings are often deemed irreproducible, or unreasonably biased, with pathogenic roles in 95% of reports. As a solution to repeatability, investigators have been told to seek funds to increase the number of human-microbiome donors (N) to increase the reproducibility of animal studies (doi:10.1016/j.cell.2019.12.025). Herein, through simulations, we illustrate that increasing N will not uniformly/universally enable the identification of consistent statistical differences (patterns of analytical irreproducibility), due to random sampling from a population with ample variability in disease and the presence of disease data subtypes (or modes). We also found that studies do not use cluster statistics when needed (97.4%, 37/38, 95%CI=86.5,99.5), and that scientists who increased N, concurrently reduced the number of mice/donor (y=-0.21x, R2=0.24; and vice versa), indicating that statistically, scientists replace the disease variance in mice by the variance of human disease. Instead of assuming that increasing N will solve reproducibility and identify clinically-predictive findings on causality, we propose the visualization of data distribution using kernel-density-violin plots (rarely used in rodent studies; 0%, 0/38, 95%CI=6.9e-18,9.1) to identify disease data subtypes to self-correct, guide and promote the personalized investigation of disease subtype mechanisms. HighlightsO_LIMultimodal diseases are those in which affected individuals can be divided into subtypes (or data modes); for instance, mild vs. severe, based on (unknown) modifiers of disease severity. C_LIO_LIThe role of the microbiome in multimodal diseases has been studied in animals; however, findings are often deemed irreproducible, or unreasonably biased, with pathogenic roles in 95% of reports. C_LIO_LIAs a solution to repeatably, investigators have been told to seek funds to increase the number of human-microbiome donors (N) to increase the reproducibility of animal studies. C_LIO_LIHerein, we illustrate that although increasing N could help identify statistical effects (patterns of analytical irreproducibility), clinically-relevant information will not always be identified. C_LIO_LIDepending on which diseases need to be compared, random sampling alone leads to reproducible patterns of analytical irreproducibility in multimodal disease simulations. C_LIO_LIInstead of solely increasing N, we illustrate how disease multimodality could be understood, visualized and used to guide the study of diseases by selecting and focusing on disease modes. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 93%
- Sex-specific transcriptome similarity networks elucidate comorbidity relationships 92%
- Gravity-based microfiltration reveals unexpected prevalence of circulating tumor cell clusters in ovarian and colorectal cancer 91%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Standardized Metric to Enhance Clinical Trial Design and Outcome Interpretation in Type 1 Diabetes 94%
- Deep representation learning for clustering longitudinal survival data from electronic health records 93%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.