Back

Prioritizing Genes and Rare Protein-Coding Variants in Acute Myeloid Leukemia via Whole Genome Sequencing Data

Vieno, S.; Singh, M.; Kramer, S.; Chatzinakos, C.; Peterson, R.; Riley, B.; Bacanu, S.-A.; Dinh, T.; Trinh, B. Q.; Nguyen, T.-H.

2026-08-22 genetic and genomic medicine
10.64898/2026.08.19.26360760 medRxiv
Show abstract

The extent to which rare and common genetic variants jointly contribute to the risk of acute myeloid leukemia (AML) still remains relatively unexplored in large-scale biobank whole-genome sequencing cohorts. Here, we leverage the latest sequencing and phenotypic data from the All of Us Research Program to identify variants, genes, and gene-sets associated with AML. We performed set-based association tests for rare protein-coding variants (Ncases=265 and Ncontrols=169,706) and single-variant association tests for common variants (Ncases=265 and Ncontrols=169,705) utilizing the large European-like ancestry sample. For the rare-variant set-based tests conducted using SAIGE-GENE+, four genes were statistically significant: DNMT3A, TET2, SRSF2, and IDH2 (Bonferroni-corrected Cauchy p-value < 0.05). We also constructed multiple rare-variant burden risk scores using different gene-sets to identify those with a substantial rare-variant burden for AML. Gene-sets derived from Genomic Data Commons whole-genome sequencing data, comprising two distinct groups-genes observed to harbor somatic mutations in AML and genes observed to harbor somatic mutations across all cancer types-showed a statistically significant rare-variant burden (Bonferroni-corrected p-value < 0.05). Ultimately, these findings demonstrate that leveraging whole-genome sequencing in large-scale biobanks enables the identification of rare protein-coding variants, genes, and gene sets associated with AML.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.