Methods for evaluating unsupervised vector representations of genomic regions
Zheng, G.; Rymuza, J.; Gharavi, E.; LeRoy, N. J.; Zhang, A.; Sheffield, N. C.
Show abstract
Representation learning models have become a mainstay of modern genomics. These models are trained to yield vector representations, or embeddings, of various biological entities, such as cells, genes, individuals, or genomic regions. Recent applications of unsupervised embedding approaches have been shown to learn relationships among genomic regions that define functional elements in a genome. Unsupervised representation learning of genomic regions is free of the supervision from curated metadata and can condense rich biological knowledge from publicly available data to region embeddings. However, there exists no method for evaluating the quality of these embeddings in the absence of metadata, making it difficult to assess the reliability of analyses based on the embeddings, and to tune model training to yield optimal results. To bridge this gap, we propose four evaluation metrics: the cluster tendency score (CTS), the reconstruction score (RCS), the genome distance scaling score (GDSS), and the neighborhood preserving score (NPS). The CTS and RCS statistically quantify how well region embeddings can be clustered and how well the embeddings preserve information in training data. The GDSS and NPS exploit the biological tendency of regions close in genomic space to have similar biological functions; they measure how much such information is captured by individual region embeddings in a set. We demonstrate the utility of these statistical and biological scores for evaluating unsupervised genomic region embeddings and provide guidelines for learning reliable embeddings. AvailabilityCode is available at https://github.com/databio/geniml
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 95%
- A Message Passing Framework for Precise Cell State Identification with scClassify2 95%
- stDyer enables spatial domain clustering with dynamic graph embedding 95%
Similar papers in this journal
- Embeddings of genomic region sets capture rich biological associations in lower dimensions 97%
- SAILER: Scalable and Accurate Invariant Representation Learning for Single-Cell ATAC-Seq Processing and Integration 95%
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 95%
- Learning interpretable representations of single-cell multi-omics data with multi-output Gaussian Processes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.