Transferable cancer detection from cell-free DNA fragment lengths through extensions of non-negative matrix factorization
Olsen, L. R.; Holsting, J. Q.; Birkbak, N. J.; Dyrskjot, L.; Pedersen, J. S.; Andersen, C. L.; Besenbacher, S.
Show abstract
MotivationThe fragment length distribution of cell-free DNA (cfDNA) can reveal the presence of circulating tumor DNA (ctDNA) in a plasma sample. A previous study documented that non-negative matrix factorization (NMF) can extract relevant features from the fragment length distributions. These distributions, however, are affected by technical biases. When NMF is performed on samples from multiple datasets, some of the extracted signatures often capture these technical biases rather than the actual biological differences. ResultsWe present two methods for extracting biologically meaningful NMF signatures across heterogeneous datasets with varying technical biases. Using simulated data, we first demonstrate that these new methods are more effective at estimating the true proportions of the underlying processes. We then show that the methods increase the transferability of fragment length signatures to cfDNA datasets from external labs. Classification models using the two proposed NMF extensions achieve an average cross-dataset AUC of 0.842 and 0.814, compared to 0.776 for standard NMF and 0.688 for a set of previously reported manually selected fragment length features. We further show that one of the proposed methods requires only 1-5 samples to estimate the batch effect during inference. Availability and ImplementationThe source code and instructions on how to run it are available at: https://github.com/BesenbacherLab/batch-NMF
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Non-negative Independent Factor Analysis disentangles discrete and continuous sources of variation in scRNA-seq data 95%
- Per-sample standardization and asymmetric winsorization lead to accurate clustering of RNA-seq expression profiles 95%
- A Bayesian framework for inter-cellular information sharing improves dscRNA-seq quantification 95%
Similar papers in this journal
- Comparison of sparse biclustering algorithms for gene expression datasets 95%
- SCDC: Bulk Gene Expression Deconvolution by Multiple Single-Cell RNA Sequencing References 94%
- Sincast: a computational framework to predict cell identities in single cell transcriptomes using bulk atlases as references 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.