Back

Interpretable gene networks from single-cell foundation models reveal conserved neurogenic dysfunction in Parkinson's disease

Yin, J.; Gosztyla, M.; Gokbag, B.; Rychkova, A.; Wilkinson, I.; St John, J.; Shah, V.; Nickels, S. L.; Schwamborn, J. C.; Fremeau, R. T.

2026-07-27 neuroscience
10.64898/2026.07.26.740813 bioRxiv
Show abstract

Interpreting large-scale single-cell transcriptomic data remains a major challenge for understanding disease mechanisms. Recent single-cell foundation models learn rich representations of gene relationships across millions of cells, yet methods for translating these embeddings into biologically interpretable gene networks remain limited. Here we present scGENet, a computational framework that constructs context-specific gene interaction networks from foundation model-derived gene embeddings. By fine-tuning pretrained models on transcriptomic data from human midbrain organoids, scGENet generates transcriptome-scale gene modules that capture biologically meaningful cellular programs. Benchmarking across multiple foundation models demonstrates that networks derived from a fine-tuned scGPT brain model show the highest concordance with curated neuronal pathways, Parkinsons disease (PD) genetic risk loci, and independent patient-derived transcriptional signatures. Applying this framework to human iPSC-derived PD midbrain organoids reveals transcriptional modules associated with neuronal differentiation, synaptic signaling, and cell-cycle regulation. Single-nucleus RNA sequencing further links these programs to altered cellular composition, including reduced dopaminergic neurons, expansion of radial glia-like progenitors, and a dopaminergic neuron subtype expressing SNCA and VGLUT2. Integration with independent human substantia nigra datasets identifies a conserved neurogenic program disrupted across genetic and idiopathic PD. Together, these results establish a generalizable strategy for extracting interpretable gene networks from single-cell foundation models, enabling systematic discovery of disease-relevant molecular programs across diverse tissues and datasets.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.