Circulating Microbial DNA as a Potential Cancer Biomarker: Technical Challenges and Controlled Evaluation
Fry Brumit, D.; Bsteh, D.; Sun, S.; Goodrich, J. A.; Alderete, T. L.; Liss, M. A.; Goldkorn, A.; Fodor, A. A.
Show abstract
Abstract Background. Circulating microbial DNA (cmDNA) has been proposed as a non-invasive cancer biomarker, but most evidence comes from cancer-sequencing datasets not designed for microbial analysis and lacking contamination controls. Whether reported signatures reflect biology or artifact is unclear in low-biomass specimens, where standard taxonomic pipelines are prone to systematic error. Methods. In a tightly controlled pilot study of metastatic castration-resistant prostate cancer, we profiled plasma cell-free DNA (cfDNA) and buffy-coat genomic DNA (gDNA) from two patients and two healthy volunteers alongside mock blood-draw and reagent controls, each with and without host-DNA depletion. Reads were classified with Kraken2/Bracken and, independently, with the marker-gene classifier MetaPhlAn. As informatics controls, reads were per-base shuffled to randomize nucleotide order while preserving read length and guanine-cytosine (GC) content, and purely synthetic reads were generated from a four-base process matched only to an aggregate GC target; both were classified identically. Genus abundances were regressed against Kraken2 database k-mer representation and against GC content. Results. Across 40 samples, Kraken2 reported several thousand genera, samples clustered by specimen type in principal-coordinate analysis (PCoA), and pooled genus counts correlated strongly with a published cancer-microbiome catalog (The Cancer Genome Atlas lung adenocarcinoma, TCGA-LUAD; Spearman {rho} = 0.81 over 282 shared genera), a pattern readily interpreted as biological signal. However, these observations were also made in per-base shuffling, which preserves GC content and length but destroys all biological sequence: shuffled reads were still abundantly classified, still clustered by specimen type, and still correlated with the catalog ({rho} {approx} 0.7), as did every sample group, including pure reagent controls. Genus counts scaled tightly with each genus's k-mer representation in the Kraken2 database on real (r 2 = 0.74) and shuffled (r2 = 0.85) reads, and the same dependence appeared in the independent published cohort. Purely synthetic reads carrying no information beyond an aggregate GC target reproduced much of the cross-cohort agreement (synthetic TCGA-LUAD {rho} = 0.61 versus 0.81 for real reads; significant in 27 of 33 TCGA cancers), and replicate shuffles of a low-GC versus a high-GC plasma sample, for which the true difference is zero, produced spurious significant differences in about 46% of genera. Regressing observed counts against the shuffled baseline left 23 genera above the artifact floor at 5% false discovery rate (FDR), nearly all known kit contaminants, control-enriched viruses, or very-low-abundance taxa; a four-criterion validity filter reduced thousands of Kraken2 genera to a single defensible candidate, Klebsiella. Conclusions. Much of the apparent cmDNA structure, including its agreement with a published cancer-microbiome catalog, is explained by base composition and reference-database architecture rather than authentic biology, and short-read k-mer pipelines cannot separate the two on their own. We find little positive evidence of an authentic circulating microbial signal, though our small sample cannot prove its absence. To limit false discovery in low-biomass metagenomics, we recommend specimen-matched negative controls, corroboration with a conservative second classifier, per-base shuffling (with GC-matched synthetic reads as a stricter floor), and GC-aware analysis.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Metagenome-assembled genomes of Estonian Microbiome cohort reveal novel species and their links with prevalent diseases 93%
- Illumina Complete Long Read Assay yields contiguous bacterial genomes from human gut metagenomes 93%
- Compositional transformations can reasonably introduce phenotype-associated values into sparse features 92%
Similar papers in this journal
- Ultra-accurate Microbial Amplicon Sequencing with Synthetic Long Reads 93%
- Application of an ecology-based analytic approach to discriminate signal and noise in low-biomass microbiome studies: whole lung tissue is the preferred sampling method for amplicon-based characterization of murine lung microbiota 92%
- Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes 92%
Similar papers in this journal
- Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding 93%
- A metagenomic DNA sequencing assay that is robust against environmental DNA contamination 92%
- Large scale capsid-mediated mobilisation of bacterial genomic DNA in the gut microbiome 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.