Back

Data Resource Profile: Genomic Data in Multiple British Birth Cohorts (1946-2001) - Health, Social, and Environmental Data from Birth to Old Age

Shireby, G.; Morris, T. T.; Wong, A.; Chaturvedi, N.; Ploubidis, G. B.; Fitzsimmons, E.; Goodman, A.; Sanchez-Galvez, A.; Davies, N. M.; Wright, L.; Bann, D.

2024-11-06 epidemiology
10.1101/2024.11.06.24316761 medRxiv
Show abstract

Birth cohort studies have a rich history of contributing to science across disciplinary fields, notably health and social sciences. Here, we introduce a curated resource comprising genomic data from five British birth cohort studies--longitudinal studies with extensive data collected prospectively across life, each deliberately sampled to be nationally representative (born 1946-2001). These contain health and social data from birth to older age, enabling longitudinal and cross-cohort genetically informed research. The Millennium Cohort Study additionally includes data on parents and offspring, enabling within-family analyses. Across five cohorts born in 1946, 1958, 1970, 1989-90, and 2000-2002, 27,432 participants have harmonized, imputed, and quality-controlled genetic data from genotyping arrays covering 6.7 million common SNPs. The Millennium Cohort Study contains over 6,000 mother-offspring pairs and over 3,000 mother-father-offspring trios. Pseudonymized data are freely available to the global research community upon approval of a data access request (https://cls.ucl.ac.uk/data-access-training).

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.