Back

A pre-trained large language model for translating single-cell transcriptome to proteome

Liu, L.; Li, W.; Wong, K.-c.; Yang, F.; Yao, J.

2023-07-04 bioinformatics
10.1101/2023.07.04.547619 bioRxiv
Show abstract

Proteins are crucial for life, and measuring their abundance at the single-cell level can facilitate a high-resolution understanding of biological mechanisms in cellular processes and disease progression. However, current single-cell proteomic technologies face challenges such as limited coverage, throughput, and sensitivity, as well as batch effects, high costs, and stringent experimental operations. Drawing inspiration from the translation procedure of both natural language processing (NLP) and the genetic central dogma, we propose a pre-trained, large generative model named scTranslator (single-cell translator). scTranslator is align-free and capable of generating multi-omics data by inferring the missing single-cell proteome based on the transcriptome. Systematic benchmarking confirms the accuracy, stability, and flexibility of scTranslator across various quantification techniques, cell types, and conditions. Furthermore, scTranslator has demonstrated its superiority in assisting various downstream analyses and applications, including gene/protein interaction inference, gene pseudo-knockout, cell clustering, batch correction, and cell origin recognition on pan-cancer data.

Published in Nature Biomedical Engineering · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.