Generative genomics accurately predicts future experimental results
Koytiger, G.; Walsh, A. M.; Marar, V.; Johnson, K. A.; Highsmith, M.; Abbas, A. R.; Stirn, A.; Brumbaugh, A. R.; David, A.; Hui, D.; Kahn, J. M.; Niu, S.-Y.; Ray, L. J.; Savonen, C.; Setvik, S.; Leek, J. T.; Bradley, R. K.
Show abstract
Realizing AIs promise to accelerate biomedical research requires AI models that are both accurate and sufficiently flexible to capture the diversity of real-life experiments. Here, we describe a generative genomics framework for AI-based experimental prediction that mirrors the process of designing and conducting an experiment in the lab or clinic. We created GEM-1 (Generate Expression Model-1), an AI system that effectively models the enormous range of bulk and single-cell gene expression experiments performed by scientists and benchmarked its performance across multiple biological axes. GEM-1s prediction of future gene expression experiments-RNA-seq data deposited in public archives after our training data cutoff-yielded accuracy comparable to the best-possible performance estimated by comparing the results of matched lab experiments. Overall, our approach illustrates the transformative potential of generative genomics for applications ranging from predicting cellular perturbations in vitro to de novo generation of data from large clinical cohorts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CellUntangler: separating distinct biological signals in single-cell data with deep generative models 97%
- Gene regulatory network inference from CRISPR perturbations in primary CD4+ T cells elucidates the genomic basis of immune disease 96%
- Variant-resolved prediction of context-specific isoform variation with a graph-based attention model 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.