ChoruMM: a versatile multi-components mixed model for bacterial-GWAS
Frouin, A.; Laporte, F.; Hafner, L.; Maury, M.; McCaw, Z. R.; Henches, L.; Julienne, H.; Chikhi, R.; Lecuit, M.; Aschard, H.
Show abstract
Genome-wide Association Studies (GWAS) have been central to studying the genetics of complex human outcomes, and there is now tremendous interest in implementing GWAS-like approaches to study pathogenic bacteria. A variety of methods have been proposed to address the complex linkage structure of bacterial genomes, however, some questions remain about to optimize the genetic modelling of bacteria to decipher causal variations from correlated ones. Here we examined the genetic structure underlying whole-genome sequencing data from 3,824 Listeria monocytogenes strains, and demonstrate that the standard human genetics model, commonly assumed by existing bacterial GWAS methods, is inadequate for studying such highly structured organisms. We leverage these results to develop ChoruMM, a robust and powerful approach that consists of a multi-component linear mixed model, where components are inferred from a hierarchical clustering of the bacteria genetic relatedness matrix. Our ChoruMM approach also includes post-processing and visualization tools that address the pervasive long-range correlation observed in bacteria genome and allow to assess the type I error rate calibration.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 96%
- Sparse haplotype-based fine-scale local ancestry inference at scale reveals recent selection on immune responses 95%
- Rapid detection of identity-by-descent tracts for mega-scale datasets 95%
Similar papers in this journal
- Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies 96%
- A resource-efficient tool for mixed model association analysis of large-scale data 95%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 95%
Similar papers in this journal
- Reconstructing the history of founder events using genome-wide patterns of allele sharing across individuals 95%
- Accurate detection of shared genetic architecture from GWAS summary statistics in the small-sample context 95%
- Regularized sequence-context mutational trees capture variation in mutation rates across the human genome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.