Back

Learning robust gene expression embeddings for disease subtyping, biomarker identification and cross-species alignment

Chen, X.; Smith, K. M.; Bi, Y.

2025-01-03 bioinformatics
10.1101/2025.01.03.631230 bioRxiv
Show abstract

High dimension gene expression measurements such as microarray and RNA-seq data, are often plagued by sources of unwanted variation. This variability can lead to the obscuring of meaningful biological signal by technical noise and non-interesting biological variation, thus resulting in failure to identify the same set of targets and biomarkers in independent studies. This phenomenon contributes to the so-called reproducibility crisis and makes preclinical drug development more challenging. There is an urgent need for an improved process to identify shared biological signals in large datasets from independent cohorts. In this article, we propose an innovative method called Joint-Embedding via Canonical Correlation Analysis (JECCA), which aims to capture relevant biological variation by learning of shared lower dimension embedding across multiple datasets. We demonstrate the efficacy of JECCA in analyzing multiple human cancer and inflammatory disease gene expression datasets, showcasing its ability to identify robust, biologically relevant signals. By leveraging these embeddings, our approach enables several important applications: (1) robust disease endotype identification by joint unsupervised analysis of multiple independent gene expression datasets (2) reproducible and biologically meaningful biomarker identification to enhance prediction of treatment outcome (3) cross-platform analysis, such as integration of RNA-seq and microarray data, which are accumulating exponentially in public databases (4) cross-species analysis to identify conserved molecular features between human and mouse. We provide comparative analyses that highlight the superior performance of JECCA over other approaches such as batch value-corrected methods based on joint analysis. As analysis of publicly available datasets becomes more prevalent, we foresee JECCA being employed in multiple disease contexts.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.