Adapting Epigenetic Clocks for Cell-Free DNA High-Throughput Sequencing Data
Li, G.; Huang, W.; Zhao, X.; Wu, J.; Guo, Y.; Chen, L.; Cao, X.; Yang, Z.; Jiang, S.; Hu, B.; Wang, Y.; Tan, D.; Tong, V.; Tang, C.; Feng, X.; Hu, X.; Ouyang, C.; Zhou, G.
Show abstract
Cell-free DNA (cfDNA) methylation sequencing holds promise for developing epigenetic aging clocks. However, current clocks--primarily trained on array-based data--do not readily generalize to high-throughput sequencing (HTS) cfDNA profiles. Using datasets with technical replicates encompassing HTS data from both cfDNA and gDNA, alongside gDNA methylation array data, we systematically assessed factors influencing clock accuracy and reproducibility. We identified key strategies to overcome HTS-specific challenges: maintaining [≥]10x mean target depth, applying elastic net regression with strong L2 regularization, and imputing unreliable beta-values. Further, transfer learning more effectively corrected platform biases than DNA type-related biases, and it enhanced performance robustly across multiple independent cohorts in aging and disease-related applications. Our findings demonstrate that array-derived epigenetic clocks can be effectively adapted to cfDNA sequencing data. This work offers critical methodological insights and practical guidelines, advancing the feasibility of minimally invasive aging assessment using cfDNA.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- SnapHiC-G: identifying long-range enhancer-promoter interactions from single-cell Hi-C data via a global background model 93%
- Accurate and fast cell marker gene identification with COSG 92%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.