Back

Adding layers of information to scRNA-seq data using pre-trained language models

Krissmer, S. M.; Menger, J.; Rollin, J.; Vogel, T. M.; Binder, H.; Hackenberg, M.

2026-03-26 bioinformatics
10.1101/2025.08.23.671699 bioRxiv
Show abstract

Pre-trained language models promise to enrich analyses of single-cell data with additional layers of information leveraging large text corpora. Yet, it is still unclear how to achieve optimal alignment with the primary quantitative single-cell data. To address this, we construct text-based training datasets from both scRNA-seq data and biomedical literature targeted to the experimental setting at hand. We then jointly train language models on both information sources to learn a common, literature-enriched representation. Our examples on functionality, disease associations, and temporal trajectories show the potential of knowledge-augmented embeddings as a generalizable and interpretable strategy for enriching single-cell analysis pipelines.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.