Batch-Harmonized Machine Learning Framework for Cross-Cohort RNA Biomarker Discovery in Pancreatic Adenocarcinoma
Markarian, M. B.; Houssigian, G. V.
Show abstract
BackgroundPancreatic ductal adenocarcinoma (PDAC) lacks reliable prognostic biomarkers. RNA-based signatures suffer from poor reproducibility due to batch effects and platform heterogeneity between microarray and RNA-seq data, limiting machine learning applications. MethodsWe developed a computational pipeline harmonizing RNA-seq data from multiple repositories using ComBat batch correction, followed by Random Forest and XGBoost classification. Restricting analysis to RNA-seq platforms only, we achieved 14,137 common genes between TCGA-PAAD (n=178) and validation cohort GSE71729 (n=357). We quantified batch correction efficacy via silhouette coefficients and trained models on survival outcomes. ResultsComBat correction eliminated dataset-specific clustering (silhouette coefficient: 0.866[->]-0.012). Random Forest achieved 64% training accuracy, identifying five prognostic biomarkers: LAMC2, DKK1, ITGB6, GPRC5A, and MAL2. These genes showed consistent importance across models and biological relevance to invasion, epithelial-mesenchymal transition, and tumor suppression. Models successfully generalized independent validation data. ConclusionsWe present the first open-source R pipeline optimized for RNA-seq-based, cross-cohort biomarker discovery in pancreatic cancer. Platform-matched datasets yielded superior gene coverage versus multi-platform approaches, enabling robust machine learning classification. Our framework identifies five novel prognostic genes and provides a reproducible method for multi-center RNA biomarker studies, available through an interactive Shiny application. AvailabilityAll code, processed data, and the interactive Shiny application are available at https://github.com/MarkBarsoumMarkarian/rna-harmonization-ai Graphical Abstract: Machine Learning Workflow for Prognostic Biomarker Discovery O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=124 SRC="FIGDIR/small/688421v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@17d6dd1org.highwire.dtl.DTLVardef@1b4caa5org.highwire.dtl.DTLVardef@643335org.highwire.dtl.DTLVardef@5dfc93_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting cancer origins with a DNA methylation-based deep neural network model 94%
- Pixelwise H-score: a novel digital image analysis based-metric to quantify membrane biomarker expression from immunohistochemistry images 94%
- Molecular signatures for inflammation vary across cancer types and correlate significantly with tumor stage, gender and vital status of patients 93%
Similar papers in this journal
- Novel ratio-metric features enable the identification of new driver genes across cancer types 94%
- Identification of patients at risk for pancreatic cancer in a 3-year timeframe based on machine learning algorithms 93%
- Deeper insights into long-term survival heterogeneity of Pancreatic Ductal Adenocarcinoma (PDAC) patients using integrative individual- and group-level transcriptome network analyses 93%
Similar papers in this journal
Similar papers in this journal
- Evaluating the Radiation Sensitivity Index and 12-chemokine gene expression signature for clinical use in a CLIA laboratory 92%
- Characterization of non-monotonic relationships between tumor mutational burden and clinical outcomes 92%
- Autofluorescence Virtual Staining System for H&E Histology and Multiplex Immunofluorescence Applied to Immuno-Oncology Biomarkers in Lung Cancer 92%
Similar papers in this journal
- driveR: A Novel Method for Prioritizing Cancer Driver Genes Using Somatic Genomics Data 94%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 92%
- Hypoxia classifier for transcriptome datasets 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.