Back

Multi-ancestry meta-analysis of tobacco use disorders based on electronic health record data prioritizes novel candidate risk genes and reveals associations with numerous health outcomes

Toikumo, S.; Jennings, M. V.; Pham, B.; Lee, H.; Mallard, T. T.; Bianchi, S. B.; Meredith, J. J.; Vilar-Ribo, L.; Xu, H.; Hatoum, A. S.; Johnson, E. C.; Pazdernik, V.; Jinwala, Z.; Leger, B. S.; Niarchou, M.; Ehinmowo, M. I.; Penn Medicine BioBank, ; MVP, ; Psychemerge Substance Use Disorder, ; Jenkins, G. D.; Batzler, A.; Pendegraft, R.; Palmer, A. A.; Zhou, H.; Biernacka, J.; Coombes, B.; Gelernter, J.; Xu, K.; Hancock, D. B.; Nancy, C. J.; Smoller, J. W.; Davis, L. K.; Justice, A. C.; Kranzler, H. R.; Kember, R. L.; Sanchez-Roige, S.

2023-03-29 genetic and genomic medicine
10.1101/2023.03.27.23287713 medRxiv
Show abstract

Tobacco use disorder (TUD) is the most prevalent substance use disorder in the world. Genetic factors influence smoking behaviors, and although strides have been made using genome-wide association studies (GWAS) to identify risk variants, the majority of variants identified have been for nicotine consumption, rather than TUD. We leveraged five biobanks to perform a multi-ancestral meta-analysis of TUD (derived via electronic health records, EHR) in 898,680 individuals (739,895 European, 114,420 African American, 44,365 Latin American). We identified 88 independent risk loci; integration with functional genomic tools uncovered 461 potential risk genes, primarily expressed in the brain. TUD was genetically correlated with smoking and psychiatric traits from traditionally ascertained cohorts, externalizing behaviors in children, and hundreds of medical outcomes, including HIV infection, heart disease, and pain. This work furthers our biological understanding of TUD and establishes EHR as a source of phenotypic information for studying the genetics of TUD.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.