Improving Differential Expression and Survival Analyses with Sample Specific Compartment Deconvolution.
Yurovsky, A.; Moffitt, R. A.
Show abstract
MotivationStudies on bulk RNA-seq of tumor biopsies can yield incorrect results because varying proportions of non-tumor tissues in the samples obscure the true signal and impact the accuracy of survival and differential expression analyses. Single-cell sequencing avoids these problems, but is still too expensive in clinical settings. Other deconvolution algorithms extract tissue-specific gene expression profiles from bulk sequencing, but cannot do this on a per-sample basis. ResultsWe introduce SSCD - sample specific compartment deconvolution. SSCD extends non-negative matrix factorization with per-sample, per-gene constraint optimization. On simulated data, SSCD shows improvements in accuracy over existing methods. Using several real cancer datasets, we show that SSCD refines Differential Expression and survival analyses. AvailabilityCode and data are available at https://github.com/ayurovsky/SSCD.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Non-negative Independent Factor Analysis disentangles discrete and continuous sources of variation in scRNA-seq data 96%
- Per-sample standardization and asymmetric winsorization lead to accurate clustering of RNA-seq expression profiles 95%
- Identification of cell-type-specific marker genes from co-expression patterns in tissue samples 95%
Similar papers in this journal
Similar papers in this journal
- MUSTANG: MUlti-sample Spatial Transcriptomics data ANalysis with cross-sample transcriptional similarity Guidance 95%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 93%
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.