scGenePT: Is language all you need for modeling single-cell perturbations?
Istrate, A.-M.; Li, D.; Karaletsos, T.
Show abstract
Modeling single-cell perturbations is a crucial task in the field of single-cell biology. Predicting the effect of up or down gene regulation or drug treatment on the gene expression profile of a cell can open avenues in understanding biological mechanisms and potentially treating disease. Most foundation models for single-cell biology learn from scRNA-seq counts, using experimental data as a modality to generate gene representations. Similarly, the scientific literature holds a plethora of information that can be used in generating gene representations using a different modality - language - as the basis. In this work, we study the effect of using both language and experimental data in modeling genes for perturbation prediction. We show that textual representations of genes provide additive and complementary value to gene representations learned from experimental data alone in predicting perturbation outcomes for single-cell data. We find that textual representations alone are not as powerful as biologically learned gene representations, but can serve as useful prior information. We show that different types of scientific knowledge represented as language induce different types of prior knowledge. For example, in the datasets we study, subcellular location helps the most for predicting the effect of single-gene perturbations, and protein information helps the most for modeling perturbation effects of interactions of combinations of genes. We validate our findings by extending the popular scGPT model, a foundation model trained on scRNA-seq counts, to incorporate language embeddings at the gene level. We start with NCBI gene card and UniProt protein summaries from the genePT approach and add gene function annotations from the Gene Ontology (GO). We name our model "scGenePT", representing the combination of ideas from these two models. Our work sheds light on the value of integrating multiple sources of knowledge in modeling single-cell data, highlighting the effect of language in enhancing biological representations learned from experimental data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CellVGAE: An unsupervised scRNA-seq analysis workflow with graph attention networks 97%
- Consensus Label Propagation with Graph Convolutional Networks for Single-Cell RNA Sequencing Cell Type Annotation 96%
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 96%
Similar papers in this journal
Similar papers in this journal
- Benchmarking imputation methods for network inference using a novel method of synthetic scRNA-seq data generation 96%
- BiGPICC: a graph-based approach to identifying carcinogenic gene combinations from mutation data 95%
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.