The Language of Cancer: Decoding Cancer Signatures with Language Models and cfRNA Sequencing
Deng, S.; Sha, L.; Jin, Y.; Zhou, T.; Wang, C.; Liu, Q.; Guo, H.; Xiong, C.; Xue, Y.; Li, X.; Li, Y.; Gao, Y.; Hong, M.; Xu, J.; Chen, S.; Wang, P.
Show abstract
We present GeneLLM, a novel large language model that offers a transformative approach to non-invasive cancer detection and biomarker discovery by directly interpreting plasma cell-free RNA (cfRNA) sequences. Unlike traditional annotation-dependent methods, GeneLLM operates without prior knowledge, achieving significantly improved multi-cancer detection accuracy. Critically, GeneLLM identifies novel cfRNAs ( pseudo-biomarkers) originating from previously unannotated genomic regions-overlooked by existing methods-offering new therapeutic targets and insights into intercellular communication. This innovative, cost-effective approach bypasses traditional bioinformatics tools, generating novel pseudo-biomarkers that outperform existing methods even with low-depth sequencing data. Consequently, GeneLLM opens new avenues for biomarker discovery and expands our understanding of the extracellular transcriptomes role in cancer development.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COSIME: Cooperative multi-view integration with Scalable and Interpretable Model Explainer 96%
- Predicting the prevalence of complex genetic diseases from individual genotype profiles using capsule networks 95%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 96%
- MORONET: Multi-omics Integration via Graph Convolutional Networks for Biomedical Data Classification 95%
- Features fusion or not: harnessing multiple pathological foundation models using Meta-Encoder for downstream tasks fine-tuning 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.