Semi-supervised segmentation and genome annotation
Chan, R. C. W.; McNeil, M.; Roberts, E. G.; Mendez, M.; Libbrecht, M. W.; Hoffman, M. M.
Show abstract
Segmentation and genome annotation methods automatically discover joint signal patterns in whole genome datasets. Previously, researchers trained these algorithms in a fully unsupervised way, with no prior knowledge of the functions of particular regions. Adding information provided by expert-created annotations to supervise training could improve the annotations created by these methods. We implemented semi-supervised learning using virtual evidence in the annotation method Segway. Additionally, we defined a positionally tolerant precision and recall metric for scoring genome annotations based on the proximity of each annotation feature to the truth set. We demonstrate semi-supervised Segways ability to learn patterns corresponding to provided transcription start sites on a specified supervision label, and subsequently recover other transcription start sites in unseen data on the same supervision label.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
- Inferring transcriptional regulators through integrative modeling ofpublic chromatin accessibility and ChIP-seq data 96%
- Harmonizing single cell 3D genome data with STARK and scNucleome 96%
Similar papers in this journal
- Embeddings of genomic region sets capture rich biological associations in lower dimensions 96%
- EpiSAFARI: Sensitive detection of valleys in epigenetic signals for enhancing annotations of functional elements 95%
- EvoAug-TF: Extending evolution-inspired data augmentations for genomic deep learning to TensorFlow 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.