A Multimodal Foundation Model for Discovering Genetic Associations with Brain Imaging Phenotypes
Machado Reyes, D.; Burch, M. C.; PARIDA, L.; Bose, A.
Show abstract
Due to the intricate etiology of neurological disorders, finding interpretable associations between multi-omics features can be challenging using standard approaches. We propose COMICAL, a contrastive learning approach leveraging multi-omics data to generate associations between genetic markers and brain imaging-derived phenotypes. COMICAL jointly learns omic representations utilizing transformer-based encoders with custom tokenizers. Our modality-agnostic approach uniquely identi-fies many-to-many associations via self-supervised learning schemes and cross-modal attention encoders. COMICAL discovered several significant associations between genetic markers and imaging-derived phenotypes for a variety of neurological disorders in the UK Biobank as well as predicting across diseases and unseen clinical outcomes from the learned representations. Source code of COMICAL along with pre-trained weights, enabling transfer learning is available at https://github.com/IBM/comical.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep feature extraction of single-cell transcriptomes by generative adversarial network 96%
- DeepPerVar: a multimodal deep learning framework for functional interpretation of genetic variants in personal genome 94%
- CLEP: A Hybrid Data- and Knowledge- Driven Framework for Generating Patient Representations 94%
Similar papers in this journal
- PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank 94%
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 94%
- SpaTM: Topic Models for Inferring Spatially Informed Transcriptional Programs 94%
Similar papers in this journal
- MORONET: Multi-omics Integration via Graph Convolutional Networks for Biomedical Data Classification 94%
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 93%
Similar papers in this journal
- Unraveling the Co-Morbidity between COVID-19 and Neurodegenerative Diseases Through Multi-scale Graph Analysis: A Systematic Investigation of Biological Databases and Text Mining 92%
- AutoGenome: An AutoML Tool for Genomic Research 90%
- Actively Protective Combinatorial Analysis: a Scalable Novel Method for Detecting Variants that Contribute to Reduced Disease Prevalence in High-Risk Individuals 90%
Similar papers in this journal
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 94%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 93%
- Meta-Analysis of the Functional Neuroimaging Literature with Probabilistic Logic Programming 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.