The Multiethnic Cohort: A Resource for the study of Genetic and non-Genetic Cancer Risk Across Populations
Bogumil, D.; Sheng, X.; Wan, P.; Xia, L.; Pooler, L.; Cheng, I.; Streicher, S.; Huang, B. Z.; Chen, F.; Stram, D.; Shen, S.; King, G.; Chiang, C. W. K.; Ongaco, C.; Adams, M.; McMullen, I.; Zhang, P.; Ling, H.; Mawhinney, M.; Doheny, K. F.; Le Marchand, L.; Wilkens, L. R.; Haiman, C. A.; Conti, D. V.
Show abstract
IntroductionThe Multiethnic Cohort Study (MEC) is a U.S. prospective cohort of over 215,000 participants, designed to investigate variation in risk factors and disease across diverse racial and ethnic groups. Over 74,000 participants contributed biospecimens for genetic studies. We describe this sub-cohort and demonstrate the types of analyses it enables. MethodsThe MEC recruited adults aged 45-75 in California and Hawaii between 1993 and 1996. Cancer diagnoses were identified via state tumor registries. The MEC Genetics Database includes 73,139 participants with germline genotype data. We evaluated genetic similarity, its relationship with self-reported race/ethnicity, and baseline characteristics, including neighborhood socioeconomic status. Using breast, colorectal, and prostate cancer as examples, the database supports multi-ancestry genome-wide association studies (GWAS), evaluation of non-genetic factors, and time-to-event analyses. ResultsParticipants included 10,962 African Americans, 24,234 Japanese Americans, 17,242 Latinos, 5,488 Native Hawaiians, 14,649 Whites, and 564 other. Principal component analysis revealed substantial diversity in ancestry. Multiethnic GWAS demonstrated effective control of population stratification while replicating many previously discovered variants. Polygenic risk score (PRS) effects varied by racial and ethnic group. Time-to-event analysis showed associations between cancer incidence and neighborhood socioeconomic status, population descriptors, and genetic similarity. DiscussionThe MEC Genetics Database enables comprehensive assessment of genetic and non-genetic cancer risk, revealing differences in absolute risk by race and ethnicity. Studying both types of risk factors in diverse and admixed populations is critical for improving risk characterization and reducing disparities. This resource supports future research in polygenic traits, gene-environment interactions, and integrated risk prediction.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Approaches for Constructing Polygenic Risk Scores for Prostate Cancer in Men of African and European Ancestry 93%
- The Phenotype-Genotype Reference Map: Improving biobank data science through replication. 93%
- The contribution of coding variants to the heritability of multiple cancer types using UK Biobank whole-exome sequencing data 93%
Similar papers in this journal
- Proteogenomic and observational evidence implicate ANGPTL4 as a potential therapeutic target for colorectal cancer prevention 92%
- Polygenic risk of any, metastatic, and fatal prostate cancer in the Million Veteran Program 92%
- A polygenic risk score for breast cancer in U.S. Latinas and Latin-American women 91%
Similar papers in this journal
- Methylation scores for smoking, alcohol consumption, and body mass index and risk of seven types of cancer 94%
- Free testosterone and malignant melanoma risk in men: prospective analyses of testosterone and SHBG with 19 cancers in men and postmenopausal women UK Biobank 93%
- Associations of plasma omega-6 and omega-3 fatty acids with overall and 19 site-specific cancers: a population-based cohort study in UK Biobank 93%
Similar papers in this journal
- Assessment of Polygenic Architecture and Risk Prediction based on Common Variants Across Fourteen Cancers 94%
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 94%
- Identifying therapeutic targets for cancer: 2,094 circulating proteins and risk of nine cancers 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.