Back

Machine learning and burden analyses highlight novel genes in Parkinson's Disease

Parlar, S. C.; Yu, E.; Kanagasingam, S.; Zhang, M.; Liu, L.; Shahkhali, M. G.; Chantereault, C.; Karpilovsky, N.; Worrall, D.; Thomas, R. A.; Senkevich, K.; Fon, E. A.; Gan-Or, Z.

2026-01-23 genetic and genomic medicine
10.64898/2026.01.22.26344646 medRxiv
Show abstract

Genome-wide association studies (GWAS) have identified numerous risk loci for Parkinsons disease, yet identifying specific causal genes remains a major challenge due to non-coding associations and complex linkage disequilibrium. Here, we present a systematic framework integrating machine learning-based gene prioritization with high-resolution rare variant burden analysis. Using an XGBoost machine-learning model trained on 285 multi-omic features, including brain-specific eQTLs and single-cell expression, we prioritized 406 out of 3,105 genes across 147 risk loci. The list of prioritized genes underwent gene- and domain-level rare variant burden analysis via optimal sequenced kernel association test (SKAT-O) across the Accelerating Medicines Partnership for Parkinsons Disease and UK Biobank cohorts (N case = 6,435, N proxy = 13,889, N control = 343,160). Meta-analysis of rare variants at the gene and domain levels replicated established associations (GBA1, LRRK2) and identified six novel potential risk genes ANKRD27 driven by p.Arg21Cys, UBXN2A, FAM171A1, ERCC8, BNC2, and ADNP. Domain-level analysis uniquely uncovered a significant association in the zinc-finger domain of ADNP, which was masked in gene-level tests. Additionally, at cohort-level, one novel gene was nominated: LRRC45 driven by p.Met607ArgfsTer57. Collectively, our results indicate that integrating machine-learning prioritization with gene- and domain-level burden testing identifies novel genes potentially involved in Parkinsons disease, with further validation needed to elucidate causality and mechanisms.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.