Phenotype harmonization and analysis for The Populations Underrepresented in Mental illness Association Studies (the PUMAS Project)
Ramirez-Diaz, A. M.; Diaz-Zuluaga, A. M.; Stroud, R. E.; Vreeker, A.; Bitta, M. A.; Ivankovic, F.; Wootton, O.; Whiteman, C. A.; Mountcastle, H.; Jha, S. C.; Georgakopoulos, P.; Kaur, I.; Mena, L.; Assaf, S.; de Souza Rodrigues, A. L.; Ziebold, C.; Newton, C. R. J. C.; Stein, D. J.; Akena, D.; Valencia-Echeverry, J.; Kyebuzibwa, J.; Palacio-Ortiz, J. D.; McMahon, J.; Ongeri, L.; Chibnik, L. B.; Quarantini, L. C.; Atwoli, L.; Santoro, M. L.; Baker, M.; Diniz, M. J. A.; Castaño-Ramirez, M.; Alemayehu, M.; Holanda, N.; Ayola-Serrano, N. C.; Lorencetti, P. G.; Mwema, R. M.; James, R.; Albuquerque,
Show abstract
BackgroundThe Populations Underrepresented in Mental illness Association Studies (PUMAS) project is attempting to remediate the historical underrepresentation of African and Latin American populations in psychiatric genetics through large-scale genetic association studies of individuals diagnosed with a serious mental illness [SMI, including schizophrenia (SCZ), schizoaffective disorder (SZA) bipolar disorder (BP), and severe major depressive disorder (MDD)] and matched controls. Given growing evidence indicating substantial symptomatic and genetic overlap between these diagnoses, we sought to enable transdiagnostic genetic analyses of PUMAS data by conducting phenotype alignment and harmonization for 89,320 participants (48,165 cases and 41,155 controls) from four cohorts, each of which used different ascertainment and assessment methods: PAISA n=9,105; PUMAS-LATAM n=14,638; NGAP n=42,953 and GPC n=22,624. As we describe here, these efforts have yielded harmonized datasets enabling us to analyze PUMAS genetic variation data at three levels: SMI overall, diagnoses, and individual symptoms. MethodsIn aligning item-level phenotypes obtained from 14 different clinical instruments, we incorporated content, branching nature, and time frame for each phenotype; standardized diagnoses; and selected 19 core SMI item-level phenotypes for analyses. The harmonization was evaluated in PUMAS cases using multiple correspondence analysis (MCA), co-occurrence analyses, and item-level endorsement. OutcomesWe mapped >6,895 item-level phenotypes in the aggregated PUMAS data, in which SCZ (44.97%) and severe BP (BP-I, 31.53%) were the most common diagnoses. Twelve of the 19 core item-level phenotypes occurred at frequencies of > 10% across all diagnoses, indicating their potential utility for transdiagnostic genetic analyses. MCA of the 14 phenotypes that were present for all cohorts revealed consistency across cohorts, and placed MDD and SCZ into separate clusters, while other diagnoses showed no significant phenotypic clustering. InterpretationOur alignment strategy effectively aggregated extensive phenotypic data obtained using diverse assessment tools. The MCA yielded dimensional scores which we will use for genetic analyses along with the item level phenotypes. After successful harmonization, residual phenotypic heterogeneity between cohorts reflects differences in branching structure of diagnostic instruments, recruitment strategies, and symptom interpretation (due to cultural variation).
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating the role of common risk variation in the recurrence risk of schizophrenia in multiplex schizophrenia families 96%
- Distinguishing clinical and genetic risk factors for suicidal ideation and behavior in a diverse hospital population 95%
- Are psychiatric disorders risk factors for COVID-19 susceptibility and severity? a two-sample, bidirectional, univariable and multivariable Mendelian Randomization study 95%
Similar papers in this journal
- Cross-phenotype relationship between opioid use disorder and suicide attempts: new evidence from polygenic association and Mendelian randomization analyses 96%
- The genomic basis of mood instability: identification of 46 loci in 363,705 UK Biobank participants, genetic correlation with psychiatric disorders, and association with gene expression and function. 95%
- Genetic heterogeneity and subtypes of major depression 95%
Similar papers in this journal
- Identifying novel subtypes of irritability using a developmental genetic approach 93%
- Reproducible Risk Loci and Psychiatric Comorbidities in Anxiety: Results from ~200,000 Million Veteran Program Participants 93%
- Penetrance and pleiotropy of polygenic risk scores for schizophrenia in 90,000 patients across three healthcare systems 93%
Similar papers in this journal
- Isolating the genetic component of mania in bipolar disorder 95%
- The genetics of the mood disorder spectrum: genome-wide association analyses of over 185,000 cases and 439,000 controls 94%
- Atypical depression is associated with a distinct clinical, neurobiological, treatment response and polygenic risk profile 94%
Similar papers in this journal
- Decoding Treatment Choice: Genetic and Phenotypic Analyses of Long-term Antidepressant Acceptability 94%
- Thinner cortex is associated with psychosis onset in individuals at Clinical High Risk for Developing Psychosis: An ENIGMA Working Group mega-analysis 94%
- Evidence for a serotonergic subtype of major depressive disorder: A NeuroPharm-1 study. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.