A multi-modal cell-free RNA language model for liquid biopsy applications
Karimzadeh, M.; Sababi, A. M.; Momen-Roknabadi, A.; Chen, N.-C.; Cavazos, T. B.; Sekhon, S.; Wang, J.; Hanna, R.; Huang, A.; Nguyen, D.; Chen, S.; Lam, T.; Hartwig, A.; Fish, L.; Li, H.; Behsaz, B.; Hormozdiari, F.; Alipanahi, B.; Goodarzi, H.
Show abstract
Cell-free RNA (cfRNA) profiling has emerged as a powerful tool for non-invasive disease detection, but its application is limited by data sparsity and complexity, especially in settings with constrained sample availability. We introduce Exai-1, a multi-modal, transformer-based generative foundation model that integrates RNA sequence embeddings with cfRNA abundance data to capture biologically meaningful representations of circulating RNAs. By leveraging both sequence and expression modalities, Exai-1 captures a biologically meaningful latent structure of cfRNA profiles. Pre-trained on over 306 billion tokens from 8,339 samples, Exai-1 enhances signal fidelity, reduces technical noise, and improves disease detection by generating synthetic cfRNA profiles. We show that self-attention and variational inference are particularly important for preservation of biological signals and contextual relationships. Additionally, Exai-1 facilitates cross-biofluid translation and assay compatibility through disentangling biological signals from confounders. By uniting sequence-informed embeddings with cfRNA expression patterns, Exai-1 establishes a transfer learning foundation for liquid biopsy, offering a scalable and adaptable framework for next-generation cfRNA-based diagnostics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 96%
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 96%
- Deep generative model embedding of single-cell RNA-Seq profiles on hyperspheres and hyperbolic spaces 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.