Back

The Language of Cancer: Decoding Cancer Signatures with Language Models and cfRNA Sequencing

Deng, S.; Sha, L.; Jin, Y.; Zhou, T.; Wang, C.; Liu, Q.; Guo, H.; Xiong, C.; Xue, Y.; Li, X.; Li, Y.; Gao, Y.; Hong, M.; Xu, J.; Chen, S.; Wang, P.

2025-03-18 bioinformatics
10.1101/2024.06.29.601341 bioRxiv
Show abstract

We present GeneLLM, a novel large language model that offers a transformative approach to non-invasive cancer detection and biomarker discovery by directly interpreting plasma cell-free RNA (cfRNA) sequences. Unlike traditional annotation-dependent methods, GeneLLM operates without prior knowledge, achieving significantly improved multi-cancer detection accuracy. Critically, GeneLLM identifies novel cfRNAs ( pseudo-biomarkers) originating from previously unannotated genomic regions-overlooked by existing methods-offering new therapeutic targets and insights into intercellular communication. This innovative, cost-effective approach bypasses traditional bioinformatics tools, generating novel pseudo-biomarkers that outperform existing methods even with low-depth sequencing data. Consequently, GeneLLM opens new avenues for biomarker discovery and expands our understanding of the extracellular transcriptomes role in cancer development.

Published in Nature Communications (predicted rank #4) · training set

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.