A large language model for predicting pancreatic ductal adenocarcinoma patients from blood-derived exosomal transcriptomics data
Choudhury, S.; Mehta, N. K.; Raghava, G. P. S.
Show abstract
Traditional machine learning approaches for text or sequence classification rely on converting textual data into numerical representations. In this study, we investigate a reverse strategy in which numerical features are transformed into sequence representations and classified using large language models (LLMs). We applied this methodology to predict pancreatic ductal adenocarcinoma (PDAC) using the expression profiles of 50 genes from 284 PDAC and 217 non-PDAC patients. Gene expression values were converted into sequence data, with each gene represented as a residue in a 50-residue protein sequence. Major LLMs like PeptideBERT, ProtBERT, and ESM2 were fine-tuned on a protein training dataset and evaluated on an independent dataset. The best-performing model, ProtBERT, achieved an AUC of 0.962 on an independent dataset. Additionally, an alignment-based approach employing BLAST and MERCI motifs was explored, and an ensemble model combining the LLM-based and alignment-based methods was developed. Our LLM-based model outperformed traditional machine learning models. To the best of our knowledge, this is the first study demonstrating the application of LLMs for mining transcriptomic profiles of cancer patients. HIGHLIGHTSO_LIIdentification of over and under-expressed genes in PDAC patients C_LIO_LIConvert numeric gene expression data to peptide sequence C_LIO_LILLM based models for predicting PDAC patients using peptide sequences C_LIO_LIMining of transcriptomics data using ProtBert and ESM2 C_LIO_LIGene expression profile for diagnostic of PDAC patients C_LI
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 95%
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 95%
- Prediction and scanning of IL-5 inducing peptides using alignment-free and alignment-based method 94%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 95%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 94%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
Similar papers in this journal
- Discovering Key Transcriptomic Regulators in Pancreatic Ductal Adenocarcinoma using Dirichlet Process Gaussian Mixture Model 94%
- In silico tool for Predicting, Designing and Scanning IL-2 inducing peptides 93%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 93%
Similar papers in this journal
- Bacteriocin Prediction Through Cross-Validation-Based and Hypergraph-Based Feature Evaluation Approaches 94%
- Predicting GD2 expression across cancer types by the integration of pathway topology and transcriptome data 93%
- BC-Predict: Mining of signal biomarkers and multilevel validation of cascade classifier for early-stage breast cancer subtyping and prognosis 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.