BulkRNABert: Cancer prognosis from bulk RNA-seq based language models
Gelard, M.; Richard, G.; Pierrot, T.; Cournede, P.-H.
Show abstract
RNA sequencing (RNA-seq) has become a key technology in precision medicine, especially for cancer prognosis. However, the high dimensionality of such data may restrict classic statistical methods, thus raising the need to learn dense representations from them. Transformers models have exhibited capacities in providing representations for long sequences and thus are well suited for transcriptomics data. In this paper, we develop a pre-trained transformer-based language model through self-supervised learning using bulk RNA-seq from both non-cancer and cancer tissues, following BERTs masking method. By probing learned embeddings from the model or using parameter-efficient fine-tuning, we then build downstream models for cancer-type classification and survival-time prediction. Leveraging the TCGA dataset, we demonstrate the performance of our method, BulkRNABert, on both tasks, with signifi-cant improvement compared to state-of-the-art methods in the pan-cancer setting for classification and survival analysis. We also show the transfer-learning capabilities of the model in the survival analysis setting on unseen cohorts. Data and Code AvailabilityIn this paper, we leverage the Cancer Genome Atlas (TCGA, https://portal.gdc.cancer.gov/), which includes bulk RNA-seq samples as well as clinical targets for each patient (cancer type, survival time). For pre-training experiments, this dataset is completed with non-cancerous bulk RNA-seq samples from GTEx (Carithers and Moore, 2015) and ENCODE (de Souza, 2012). Code available at https://github.com/instadeepai/multiomics-open-research Institutional Review Board (IRB)This research does not require IRB approval.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 94%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
- Accessible, Reproducible, and Scalable Machine Learning for Biomedicine 94%
Similar papers in this journal
- A Network-centric Framework for theEvaluation of Mutual Exclusivity Tests onCancer Drivers 94%
- Automatic recognition of complementary strands: Lessons regarding machine learning abilities in RNA folding 94%
- Machine Learning Approaches Identify Genes Containing Spatial Information from Single-Cell Transcriptomics Data. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.