Phenotypic subtyping via contrastive learning
Gorla, A.; Sankararaman, S.; Burchard, E.; Flint, J.; Zaitlen, N.; Rahmani, E.
Show abstract
Defining and accounting for subphenotypic structure has the potential to increase statistical power and provide a deeper understanding of the heterogeneity in the molecular basis of complex disease. Existing phenotype subtyping methods primarily rely on clinically observed heterogeneity or metadata clustering. However, they generally tend to capture the dominant sources of variation in the data, which often originate from variation that is not descriptive of the mechanistic heterogeneity of the phenotype of interest; in fact, such dominant sources of variation, such as population structure or technical variation, are, in general, expected to be independent of subphenotypic structure. We instead aim to find a subspace with signal that is unique to a group of samples for which we believe that subphenotypic variation exists (e.g., cases of a disease). To that end, we introduce Phenotype Aware Components Analysis (PACA), a contrastive learning approach leveraging canonical correlation analysis to robustly capture weak sources of subphenotypic variation. In the context of disease, PACA learns a gradient of variation unique to cases in a given dataset, while leveraging control samples for accounting for variation and imbalances of biological and technical confounders between cases and controls. We evaluated PACA using an extensive simulation study, as well as on various subtyping tasks using genotypes, transcriptomics, and DNA methylation data. Our results provide multiple strong evidence that PACA allows us to robustly capture weak unknown variation of interest while being calibrated and well-powered, far superseding the performance of alternative methods. This renders PACA as a state-of-the-art tool for defining de novo subtypes that are more likely to reflect molecular heterogeneity, especially in challenging cases where the phenotypic heterogeneity may be masked by a myriad of strong unrelated effects in the data. Code AvailabilityPACA is available as an open source R package on GitHub: https://github.com/Adigorla/PACA
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 94%
- A spectral framework to map QTLs affecting joint differential networks of gene co-expression 94%
- Cross-GWAS coherence test at the gene and pathway level 94%
Similar papers in this journal
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 95%
- Co-expression-wide association studies link genetically regulated interactions with complex traits 95%
- Fast Kernel-based Association Testing of non-linear genetic effects for Biobank-scale data 94%
Similar papers in this journal
- Fast and Accurate Bayesian Polygenic Risk Modeling with Variational Inference 95%
- Sparse modeling of interactions enables fast detection of genome-wide epistasis in biobank-scale studies 95%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 95%
Similar papers in this journal
- Generative prediction of causal gene sets responsible for complex traits 95%
- Mendelian Randomization for causal inference accounting for pleiotropy and sample structure using genome-wide summary statistics 95%
- Dissecting heterogeneous cell-populations across drug and disease conditions with PopAlign 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.