FedPyDESeq2: a federated framework for bulk RNA-seq differential expression analysis
Muzellec, B.; Marteau-Ferey, U.; Marchand, T.
Show abstract
Large-scale transcriptomic studies are often limited by data silos and risks of privacy leakage, which may lead to missed clinical insights. Meta-analysis methods may be used to aggregate local results, but they induce lower statistical power and are particularly sensitive to heterogeneous settings. A recent paradigm in distributed computing, federated learning (FL) is a means of fitting models from siloed data, while ensuring that private data does not leave its storage facilities. Here, we introduce FedPyDESeq2, a software for differential expression analysis (DEA) on siloed bulk RNA-seq. Building on FL tools, FedPyDESeq2 implements the DESeq2 pipeline for DEA on siloed datasets in a privacy-enhancing manner. We benchmark FedPyDESeq2 on datasets from The Cancer Genome Atlas corresponding to 8 different indications, split by geographical origin. FedPyDESeq2 achieves near-identical results on siloed data compared with PyDESeq2 on pooled data, and significantly outperforms meta-analysis baselines.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 95%
- BANDITS: Bayesian differential splicing accounting for sample-to-sample variability and mapping uncertainty 95%
- PseudotimeDE: inference of differential gene expression along cell pseudotime with well-calibrated p-values from single-cell RNA sequencing data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.