De novo characterization of cell-free DNA fragmentation hotspots boosts the power for early detection and localization of multi-cancer
Zhou, X.; Liu, Y.
Show abstract
Non-random cell-free DNA fragmentation is a promising signature for cancer diagnosis. However, its aberration at the fine-scale in early-stage cancers is poorly understood. Here, we developed an approach to de novo characterize the cell-free DNA fragmentation hotspots from whole-genome sequencing. In healthy, hotspots are enriched in gene-regulatory elements, including open chromatin regions, promoters, hematopoietic-specific enhancers, and, interestingly, 3end of transposons. Hotspots identified in early-stage hepatocellular carcinoma patients showed overall hypo-fragmentation patterns compared to healthy controls. These cancer-specific hypo-fragmented hotspots are associated with genes enriched in gene ontologies and KEGG pathways that are related to the initiations of hepatocellular carcinoma and cancer stem cells. Further, we identified the fragmentation hotspots at 297 cancer samples across 8 different cancer types (92% in stage I to III), 103 benign samples, and 247 healthy samples. The fine-scale fragmentation level at most variable hotspots showed cancer-specific fragmentation patterns across multiple cancer types and non-cancer controls. With the fine-scale fragmentation signals alone in a machine learning model, we achieved 48% to 95% sensitivity at 100% specificity in different early-stage cancer. We further validated the model at independent datasets we generated at a small number of early-stage cancers and healthy plasma samples with matched age, gender, and lifestyle. In cancer-positive cases, we further localized cancer to a small number of anatomic sites with a median of 80% accuracy. The results highlight the significance of de novo characterizing the cell-free DNA fragmentation hotspots for detecting early-stage cancers and dissection of gene-regulatory aberrations in cancers.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FinaleMe: Predicting DNA methylation by the fragmentation patterns of plasma cell-free DNA 97%
- DNA 5-methylcytosine detection and methylation phasing using PacBio circular consensus sequencing 96%
- Cross-dataset pan-cancer detection: Correlating cell-free DNA fragment coverage with open chromatin sites across cell types 96%
Similar papers in this journal
- Chromatin Interaction Neural Network (ChINN): A machine learning-based method for predicting chromatin interactions from DNA sequences 96%
- Comprehensive characterization of single cell full-length isoforms in human and mouse with long-read sequencing 95%
- SOAPy: a Python package to dissect spatial architecture, dynamics and communication 95%
Similar papers in this journal
- Multi-sample Full-length Transcriptome Analysis of 22 Breast Cancer Clinical Specimens with Long-Read Sequencing 96%
- Comprehensive multimodal and multiomic profiling reveals epigenetic and transcriptional reprogramming in lung tumors 95%
- Gene Function Revealed at the Moment of Stochastic Gene Silencing 95%
Similar papers in this journal
- Polygenic regression uncovers trait-relevant cellular contexts through pathway activation transformation of single-cell RNA sequencing data 94%
- Trans-eQTL mapping in gene sets identifies network effects of genetic variants 94%
- Scalable Screening of Ternary-Code DNA methylation Dynamics Associated with Human Traits. 94%
Similar papers in this journal
- Fast Fourier Transform is a training-free, ultrafast, highly efficient, and fully interpretable approach for epigenomic data compression 96%
- Z-Flipons conserved between human and mouse are associated with increased transcription initiation rates 96%
- Integrative Network Analysis of Differentially Methylated and Expressed Genes for Biomarker Identification in Leukemia 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.