Semi-supervised learning with pseudo-labeling for regulatory sequence prediction
Mourad, R.
Show abstract
Predicting molecular processes using deep learning is a promising approach to provide biological insights for non-coding SNPs identified in genome-wide association studies. However, most deep learning methods rely on supervised learning, which requires DNA sequences associated with functional data, and whose amount is severely limited by the finite size of the human genome. Conversely, the amount of mammalian DNA sequences is growing exponentially due to ongoing large-scale sequencing projects, but in most cases without functional data. To alleviate the limitations of supervised learning, we propose a novel semi-supervised learning based on pseudo-labeling, which allows to exploit unlabeled DNA sequences from numerous genomes during model pre-training. The approach is very flexible and can be used to train any neural architecture including state-of-the-art models, and shows in certain situations strong predictive performance improvements compared to standard supervised learning in most cases. Moreover, small models trained by SSL showed similar or better performance than large language model DNABERT2.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 95%
- Pair consensus decoding improves accuracy of neural network basecallers for nanopore sequencing 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.