Back

Data Preparation of the nuMoM2b Dataset

Goretsky, A.; Dmitrienko, A.; Tang, I.; Lari, N.; Kunhardt, O.; Rashid Khan, R.; Marcussen, C. E.; Catto, A.; Mallia, D.; Leshchenko, A.; Lin, A.; Raja, A.; Salleb-Aouissi, A.; Pe'er, I.; Wapner, R.; Gyamfi-Bannerman, C.

2021-08-26 obstetrics and gynecology
10.1101/2021.08.24.21262142 medRxiv
Show abstract

In 2010, the Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD) started the Nulliparous Pregnancy Outcomes Study: Monitoring Mothers-to-be (nuMoM2b), a prospective cohort study of a racially/ethnically/geographically diverse population of nulliparous women with singleton gestation. The nuMoM2b is a very large dataset, consisting of data for 10,038 patients with over 4,600 features per patient, spread out over 80 files. In this report, we share our experience preparing and working with this dataset. We present our data preprocessing of the nuMoM2b dataset to get a deeper understanding of the data, extract the most relevant features, make the fewest assumptions when filling in unknown values, and reducing the dimensionality of the data. We hope this report is useful to researchers interested in building machine learning and statistical models from the nuMoM2b dataset.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.