Adding layers of information to scRNA-seq data using pre-trained language models
Krissmer, S. M.; Menger, J.; Rollin, J.; Vogel, T. M.; Binder, H.; Hackenberg, M.
Show abstract
Pre-trained language models promise to enrich analyses of single-cell data with additional layers of information leveraging large text corpora. Yet, it is still unclear how to achieve optimal alignment with the primary quantitative single-cell data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then jointly train language models on both information sources to learn a common, literature-enriched representation. Our examples on functionality, disease associations, and temporal trajectories show the potential of knowledge-augmented embeddings as a generalizable and interpretable strategy for enriching single-cell analysis pipelines.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 97%
- Learning interpretable representations of single-cell multi-omics data with multi-output Gaussian Processes 96%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 95%
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 97%
- multiDGD: A versatile deep generative model for multi-omics data 96%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.