Back

sc-pcQTL: hurdle-based co-expression modeling for multi-gene QTL mapping in single-cell RNA-seq data

Zhang, J.; Huang, Y.; Claussnitzer, M.; Kanai, M.; Zhou, W.

2026-08-23 bioinformatics
10.64898/2026.08.18.745314 bioRxiv
Show abstract

Motivation: Single-cell expression quantitative trait locus (eQTL) studies can resolve cell-type-specific genetic effects, but conventional gene-by-gene analyses do not directly capture coordinated genetic regulation of neighboring genes. Principal-component QTL (pcQTL) mapping can summarize such multi-gene effects, but existing approaches were developed for bulk expression and are not designed for sparse single-cell counts. Results: We developed sc-pcQTL, a framework that applies two-component hurdle modeling and sliding-window clustering to identify local co-expression clusters, summarizes each cluster using principal components, and maps cis-pcQTLs. In simulations, the individual hurdle components controlled type I error, while the component-union screening rule was substantially more powerful than donor-level pseudobulk correlation tests. Applied to 1.24 million peripheral blood mononuclear cells from 982 OneK1K donors across 10 cell types, sc-pcQTL identified 2,485 local co-expression clusters and conducted QTL mapping for 4,353 cluster-PC phenotypes at single-cell resolution, of which 2,040 had at least one significant cis-pcQTL association. Fine-mapping and colocalization with genome-wide association study loci across 1,163 phenotypes in the FinnGen study identified 394 colocalized QTL-GWAS signal groups. Each group comprised fine-mapped QTL and GWAS signals connected through one or more colocalization links within the same cell type and local gene cluster. Of these groups, 46 were pcQTL-specific and contained no colocalized single-gene eQTL from a constituent gene. Locus-level analyses further revealed cell-type-specific multi-gene regulatory effects. Thus, sc-pcQTL complements conventional single-gene eQTL analysis by identifying trait-relevant regulatory signals shared across neighboring genes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.