Back

cellSTAAR: Incorporating single-cell-sequencing-based functional data to boost power in rare variant association testing of non-coding regions

Van Buren, E.; Zhang, Y.; Li, X.; Selvaraj, M. S.; Li, Z.; Zhou, H.; Palmer, N. D.; Arnett, D. K.; Blangero, J.; Boerwinkle, E.; Cade, B. E.; Carlson, J. C.; Carson, A. P.; Chen, Y.-D. I.; Curran, J.; Duggirala, R.; Fornage, M.; Franceschini, N.; Graff, M.; Gu, C.; Guo, X.; He, J.; Heard-Cosa, N.; Hou, L.; Hung, Y.-J.; Kalyani, R. R.; Kardia, S. L. R.; Kooperberg, C.; Kral, B. G.; Lange, L.; Li, C.; Liu, S.; Lloyd-Jones, D.; Loos, R. J. F.; Manichaikul, A. W.; Martin, L. W.; Mathias, R.; Minster, R.; Mitchell, B. D.; Mychaleckyj, J. C.; Naseri, T.; North, K.; O'Connell, J.; Perry, J. A.; Peyse

2025-04-26 genetics
10.1101/2025.04.23.650307 bioRxiv
Show abstract

Whole genome sequencing (WGS) studies have identified hundreds of millions of rare variants (RVs) and have enabled RV association tests (RVATs) of these variants with complex traits and diseases. Analysis of non-coding variants is challenged by the considerable variability in regulatory function which candidate Cis-Regulatory Elements (cCREs) exhibit across cell types. We propose cellSTAAR, which integrates WGS data with single-cell ATAC-seq data to capture variability in chromatin accessibility across cell types via the construction of cell-type-specific functional annotations and variant sets. To reflect the uncertainty in cCRE-gene linking, cellSTAAR also links cCREs to their target genes using an omnibus framework which aggregates results from a variety of popular linking approaches. We applied cellSTAAR on Freeze 8 (N = 60,000) of the NHLBI Trans-Omics for Precision Medicine (TOPMed) consortium data to four lipids phenotypes: LDL cholesterol, a binary variable corresponding to high LDL cholesterol, HDL cholesterol, and triglycerides. We also provide replication results for all four phenotypes using UK Biobank (N = 190,000). Evidence from simulation studies and our real data analysis demonstrates that cellSTAAR boosts power and improves interpretation of RVATs of cCREs.

Published in Nature Methods · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.