Learning a latent representation of human genomics using Avocado
Schreiber, J.; Noble, W. S.
Show abstract
In the past decade, the use of high-throughput sequencing assays has allowed researchers to experimentally acquire thousands of functional measurements for each basepair in the human genome. Despite their value, these measurements are only a small fraction of the potential experiments that could be performed while also being too numerous to easily visualize or compute on. In a recent pair of publications, we address both of these challenges with a deep neural network tensor factorization method, Avocado, that compresses these measurements into dense, information-rich representations. We demonstrate that these learned representations can be used to impute with high accuracy the output of experimental assays that have not yet been performed and that machine learning models that leverage these representations outperform those trained directly on the functional measurements on a variety of genomics tasks. The code is publicly available at https://github.com/jmschrei/avocado.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
- Ultrafast and interpretable single-cell 3D genome analysis with Fast-Higashi 94%
- Geometric Sketching Compactly Summarizes the Single-Cell Transcriptomic Landscape 94%
Similar papers in this journal
- AdaLiftOver: High-resolution identification of orthologous regulatory elements with adaptive liftOver 95%
- SAILER: Scalable and Accurate Invariant Representation Learning for Single-Cell ATAC-Seq Processing and Integration 95%
- DECODE: A Deep-learning Framework for Condensing Enhancers and Refining Boundaries with Large-scale Functional Assays 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.