Back

Whole-Genome Promoter Profiling of Plasma Cell-Free DNA Exhibits Predictive Value for Preterm Birth

Guo, Z.-W.; Wang, K.; Huang, X.; Li, K.; Ouyang, G.; Yang, X.; Tan, J.; Shi, H.; Luo, L.; Zhang, X.; Zhang, M.; Han, B.-W.; Zhai, X.; Wu, Y.; Yang, F.; Yang, X.-X.; Tang, J.

2022-09-23 genetic and genomic medicine
10.1101/2022.09.20.22280143 medRxiv
Show abstract

Preterm birth (PTB) occurs in around 11% of all births worldwide, resulting in significant morbidity and mortality for both mothers and offspring. Identification of pregnancies at risk of preterm birth in early pregnancy may help improve intervention and reduce its incidence. However, there exist few methods for PTB prediction developed with large sample size, high throughput screening and validation in independent cohorts. Here, we established a large-scale, multi-center, and case-control study that included 2,590 pregnancies (2,072 full-term and 518 preterm pregnancies) from three independent hospitals to develop a preterm birth classifier. We implemented whole-genome sequencing on their plasma cfDNA and then their promoter profiling (read depth spanning from -1 KB to +1 KB around the transcriptional start site) was analyzed. Using three machine learning models and two feature selection algorithms, classifiers for predicting preterm delivery were developed. Among them, a classifier based on the support vector machine model and backward algorithm, named PTerm (Promoter profiling classifier for preterm prediction), exhibited the largest AUC value of 0.878 (0.852-0.904) following LOOCV cross-validation. More importantly, PTerm exhibited good performance in three independent validation cohorts and achieved an overall AUC of 0.849 (0.831-0.866). Taken together, PTerm could be based on current noninvasive prenatal test (NIPT) data without changing its procedure or adding detection cost, which can be easily adapted for preclinical tests.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.