Back

Monte Carlo Simulation Analysis of Variant Burden Prompts Potential Oligogenic Interactions

Pata, V.; Marandi, M.; Huik, J. M.; Ounap, K.; Pajusalu, S.

2025-12-27 bioinformatics
10.64898/2025.12.27.696666 bioRxiv
Show abstract

BackgroundA large fraction of patients with suspected rare genetic disorders remain undiagnosed, suggesting that complex inheritance patterns, which are poorly detected by standard monogenic tests, may be a key factor. MethodsWe developed a computational framework that uses Monte Carlo simulations on summary-level allele counts to test for pathway-level burden enrichment. This approach was applied to a cohort of 639 neuromuscular disease (NMD) patients and 9,059 controls from the Tartu University Hospital (2015-2023) as proof of concept and compared with standard burden testing. ResultsWhile standard per-gene Z-tests failed to yield genome-wide significant results, our simulation framework identified a nominal enrichment of rare, moderate-impact variants within a predefined set of NMD-associated genes (Bonferroni-corrected empirical p = 0.02). Exploratory analysis of the most significantly burdened gene sets revealed frequent co-occurrence of CAPN3 with other critical NMD genes, including CLCN1, GFPT1, and MYOT. ConclusionOur findings suggest a potential oligogenic model in NMD, where variants in a hub gene like CAPN3 may interact with other rare variants to contribute to disease. This Monte Carlo simulation framework needs further validation but demonstrates that this framework may be a useful tool for hypothesis generation in complex genetic disorders. FundingEstonian Research Council grants PSG774 and PRG2040. Availabilitypublicly available under CC BY-NC-SA 4.0 licence at https://github.com/OligoGeneticDiseases/gen-toolbox

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.