Back

A large language model for predicting pancreatic ductal adenocarcinoma patients from blood-derived exosomal transcriptomics data

Choudhury, S.; Mehta, N. K.; Raghava, G. P. S.

2025-03-11 bioinformatics
10.1101/2025.03.06.641795 bioRxiv
Show abstract

Traditional machine learning approaches for text or sequence classification rely on converting textual data into numerical representations. In this study, we investigate a reverse strategy in which numerical features are transformed into sequence representations and classified using large language models (LLMs). We applied this methodology to predict pancreatic ductal adenocarcinoma (PDAC) using the expression profiles of 50 genes from 284 PDAC and 217 non-PDAC patients. Gene expression values were converted into sequence data, with each gene represented as a residue in a 50-residue protein sequence. Major LLMs like PeptideBERT, ProtBERT, and ESM2 were fine-tuned on a protein training dataset and evaluated on an independent dataset. The best-performing model, ProtBERT, achieved an AUC of 0.962 on an independent dataset. Additionally, an alignment-based approach employing BLAST and MERCI motifs was explored, and an ensemble model combining the LLM-based and alignment-based methods was developed. Our LLM-based model outperformed traditional machine learning models. To the best of our knowledge, this is the first study demonstrating the application of LLMs for mining transcriptomic profiles of cancer patients. HIGHLIGHTSO_LIIdentification of over and under-expressed genes in PDAC patients C_LIO_LIConvert numeric gene expression data to peptide sequence C_LIO_LILLM based models for predicting PDAC patients using peptide sequences C_LIO_LIMining of transcriptomics data using ProtBert and ESM2 C_LIO_LIGene expression profile for diagnostic of PDAC patients C_LI

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.