Mendelianization: Concentrating Polygenic Signal into a Single Causal Locus
Strobl, E. V.
Show abstract
Complex disorders such as depression and alcohol use involve numerous genetic variants, and implicated loci continue to grow with sample size. This proliferation hampers interpretability, as the mechanisms by which so many variants jointly contribute to pathophysiology remain unclear. In contrast, classical Mendelian diseases arise from a single causal locus and are easier to interpret. We thus introduce Mendelianization - an algorithm distinct from Mendelian randomization - that learns weighted combinations of outcomes so that each aggregated phenotype concentrates association at one locus. We prove that this locus is causal under four structural assumptions natural to genetic data. The method handles partial sample overlap, provides calibrated hypothesis tests, maps coefficients to interpretable scales, and quantifies the degree of Mendelianism using summary z-statistics alone. In experiments, Mendelianization enhances statistical power to detect Mendelian symptom profiles even in heterogeneous disorders like major depression, generalized anxiety, and alcohol use disorder. An R implementation is available at github.com/ericstrobl/Mendelianization.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. 95%
- Quantifying genetic effects on disease mediated by assayed gene expression levels 95%
- Fast and flexible joint fine-mapping of multiple traits via the Sum of Single Effects model 95%
Similar papers in this journal
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 96%
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 96%
- Efficient variance components analysis across millions of genomes 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.