Semi-supervised deep learning with graph neural network for cross-species regulatory sequence prediction
Mourad, R.
Show abstract
Genome-wide association studies have systematically identified thousands of single nucleotide polymorphisms (SNPs) associated with complex genetic diseases. However, the majority of those SNPs were found in non-coding genomic regions, preventing the understanding of the underlying causal mechanism. Predicting molecular processes based on the DNA sequence represents a promising approach to understand the role of those non-coding SNPs. Over the past years, deep learning was successfully applied to regulatory sequence prediction. Such method required DNA sequences associated with functional data for training. However, the human genome has a finite size which strongly limits the amount of DNA sequence with functional data available for training. Conversely, the amount of mammalian DNA sequences is exponentially increasing due to ongoing large sequencing projects, but without functional data in most cases. Here, we propose a semi-supervised learning approach based on graph neural network which allows to borrow information from homologous mammal sequences during training. Our approach can be plugged into any existing deep learning model and showed improvements in many different situations, including classification and regression, and for different types of functional data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepPHiC: Predicting promoter-centered chromatin interactions using a novel deep learning approach 97%
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 96%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 96%
Similar papers in this journal
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 96%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
- Biologically-relevant transfer learning improves transcription factor binding prediction 96%
Similar papers in this journal
- A Comprehensive Evaluation of Self Attention for Detecting Regulatory Feature Interactions 97%
- Integrating Protein and DNA Embeddings for Improving Genome-Wide Transcription Factor Binding Site Prediction 96%
- Transfer Learning Compensates Limited Data, Batch-Effects, And Technical Heterogeneity In Single-Cell Sequencing 96%
Similar papers in this journal
- NIMBus: a Negative Binomial Regression based Integrative Method for Mutation Burden Analysis 96%
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 95%
- Single-cell Multi-omics Integration for Unpaired Data by a Siamese Network with Graph-based Contrastive Loss 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.