Back

Identification of cell types associated with 14 brain phenotypes from more than 10 million single cells

Phung, T. N.; Seoane, S. L.; Li, W.-P.; Brouwer, R.; Posthuma, D. N.

2025-12-09 genetics
10.64898/2025.12.05.692533 bioRxiv
Show abstract

Genome-wide association studies (GWAS) have yielded unprecedented insight into the genetic variants associated with many human traits. However, translating this insight into knowledge about causal biological mechanisms remains challenging. One promising recent strategy is to link GWAS-associated genes to information on expression levels of those genes in specific cell types. This strategy allows for the generation of specific, testable hypotheses about which cell types are important for a trait which can then be investigated in functional lab experiments for actual relevance. The success of this strategy strongly depends on the quality and systematic analysis of available single-cell RNAseq datasets. Here we present a comprehensive database of 388 datasets derived from 36 studies spanning different regions of the brain across developmental timepoints. Using this database, we tested for the presence of cell type enrichment in genes associated with 14 traits. We confirmed previous findings such as the association between microglia and Alzheimers disease in the entorhinal cortex. In addition, we found novel evidence for the involvement of specific cell types in disease, such as astrocytes being implicated in alcohol-related phenotypes or neuronal cell types at the prenatal stage in ADHD. Our database has been incorporated to FUMA, a publicly available and widely used tool for post-GWAS functional annotation analyses. Our work provides an approach to facilitate the prioritization of cell types in specific brain regions and/or developmental stages, further informing the design of follow-up experiments.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.