High Precision Binary Trait Association on Phylogenetic Trees
Balogun, I. O.; Mancuso, C. P.; Lieberman, T. D.
Show abstract
Traditional methods for identifying associations between genomic features and traits, or between pairs of genomic traits, struggle when applied to bacterial genomes. While several microbial GWAS (mGWAS) methods have been developed to account for the fact that genome-wide linkage in bacteria creates strong evolutionary-induced associations, these methods have high false discovery rates or lack statistical power, have poor performance on negative interactions, and face computational limits at the scale required for pangenome-wide study of gene-gene interactions. Here, we present SimPhyNI, a computationally optimized framework for efficient and rigorous mGWAS studies. SimPhyNI builds null co-occurrence distributions by independently simulating traits using phylogenetically-informed parameters, novelly including time to first event. The constrained variation in these simulations, combined with log odds ratio scoring for comparing across traits, robustly identifies both positive and negative associations. Using synthetic datasets mimicking both gene-gene and gene-trait associations, we demonstrate that SimPhyNI achieves high precision and recall for both positive and negative interactions. We demonstrate SimPhyNIs utility by detecting interactions between phage defense systems in E. coli and gene-gene interactions across the entire E. coli pangenome (>9 million tests). Though developed here for binary traits, SimPhyNIs design supports extension to multi-state and continuous traits using generalized models of stochastic simulation. SimPhyNIs performance and scalability enable genome-wide discovery of genetic interactions that drive microbial function, ecology, and disease. Impact StatementUnderstanding how bacterial genes associate with traits and with one another is essential for predicting disease outcomes, antibiotic resistance, and future evolution. However, identifying these interactions is challenging because shared ancestry creates false correlations. SimPhyNI overcomes this through an ancestry-informed statistical simulation process, achieving near-zero false positive rates while maintaining computational efficiency for large scale analyses. This efficiency enables systematic mapping of gene-gene interaction networks across large datasets containing thousands of genes and genomes. As microbial genomic datasets continue to expand, SimPhyNIs scalability and precision will accelerate discovery of the mechanistic principles underlying infectious disease, microbiome function, and microbial evolution and ecology.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Machine-learning predicts genomic determinants of meiosis-driven structural variation in a eukaryotic pathogen 94%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 94%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 94%
Similar papers in this journal
- Estimating maximal microbial growth rates from cultures, metagenomes, and single cells via codon usage patterns 95%
- Phase variation as a major mechanism of adaptation in Mycobacterium tuberculosis complex 93%
- Genomic islands of differentiation in a rapid avian radiation have been driven by recent selective sweeps 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.