Back

Sequencing and health data resource of children of African ancestry

Kottyan, L.; Richards, S.; Tracy, M.; Lawson, L. P.; Cobb, B. L.; Esslinger, S.; Gerwe, M.; Morgan, J.; Chandel, A.; Travitz, L.; Huang, Y.; Black, C.; Sobowale, A.; Akintobi, T.; Mitchell, M.; Beck, A. F.; Unaka, N.; Seid, M.; Fairbanks, S.; Adams, M.; Mersha, T.; Namjou, B.; Pauciulo, M. W.; Strawn, J. R.; Ammerman, R. T.; Santel, D.; Pestian, J.; Glauser, T.; Prows, C.; Martin, L. J.; Muglia, L.; Harley, J. B.; Chepelev, I.; Kaufman, K. M.

2025-03-26 genetic and genomic medicine
10.1101/2025.03.22.25324419 medRxiv
Show abstract

PurposeIndividuals who self-report as Black or African American are historically underrepresented in genome-wide studies of disease risk, a disparity particularly evident in pediatric disease research. To address this gap, Cincinnati Childrens Hospital Medical Center (CCHMC) established a biorepository and developed a comprehensive DNA sequencing resource including 15,684 individuals who self-identified as African American or Black and received care at CCHMC. MethodsParticipants were enrolled through the CCHMC Discover Together Biobank and sequenced. Admixture analyses confirmed the genetic ancestry of the cohort, which was then linked to electronic medical records. ResultsHigh-quality genome-wide genotypes from common variants accompanied by medical record-sourced data are available through the Genomic Information Commons. This dataset performs well in genetic studies. Specifically, we replicated known associations in sickle cell disease (HBB, HGNC:4827, p = 4.05 x 10-{superscript 1}LL), anxiety (PLAAT3, HGNC:17825, p = 6.93 x 10-L), and asthma (PCDH15, HGNC:14674, p = 5.6 x 10-{superscript 1}L), while also identifying novel loci associated with anxiety, asthma, and asthma severity. ConclusionWe present the acquisition and quality of genetic and disease-associated data and present an analytical framework for using this resource. In partnership with a community advisory council, we have co-developed a valuable framework for data use and future research.

Published in Genetics in Medicine (predicted rank #8) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.