Endogenous labeling empowers accurate detection of m6A from single long reads of direct RNA sequencing
Guo, W.; Ren, Z.; Huang, X.; He, J.; Zhang, J.; Wu, Z.; Guo, Y.; Zhang, Z.; Cun, Y.; Wang, J.
Show abstract
Although plenty of machine learning models have been developed to detect m6A RNA modification sites using the electric current signals of ONT direct RNA sequencing (DRS) reads, the landscape of m6A on different RNA isoforms is still a mystery due to their limited capacity to distinguish the m6A on individual long reads and RNA isoforms. The primary challenge in training the model with single-read accuracy is the difficulty of obtaining the training data from individual DRS reads that comprehensively represent the m6A on endogenous RNAs. Here, we endogenously label the methylated m6A sites on single ONT DRS reads by APOBEC1-YTH induced C-to-U mutations, strategically positioned 10-100 nt away from the known m6A sites on the same reads. Adopting a semi-supervised leaning strategy, we obtain 700,438 reliable 5-mer single-read level m6A signals, providing a comprehensive representation of m6A on endogenous RNAs. Leveraging this dataset, we develop m6Aiso, a deep residual neural network model that not only accurately identifies and quantifies known m6A sites but also reveals unknown, subtly methylated m6A sites responsive to METTL3 depletion. Analyzing m6Aiso-determined m6A on single reads and isoforms uncovers distance-dependent linkages of m6A sites along single molecules, as well as differential methylation of identical m6A sites on different isoforms. Moreover, we find wide-spread functionally important dynamic changes of m6A sites on specific isoforms during epithelial-mesenchymal transition (EMT). The pivotal utilization of the endogenous labeling strategy empowers m6Aiso to achieve remarkable precision in pinpointing m6A on individual molecules, underscores its effectiveness in elucidating the intricate dynamics and complexities of m6A across RNA isoforms.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 97%
- CapTrap-Seq: A platform-agnostic and quantitative approach for high-fidelity full-length RNA transcript sequencing 96%
- MePMe-seq: Antibody-free simultaneous m6A and m5C mapping in mRNA by metabolic propargyl labeling and sequencing 96%
Similar papers in this journal
- Functional classification of noncoding RNAs associated with distinct histone modifications by PIRCh-seq 97%
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 96%
- Comprehensive analyses of partially methylated domains and differentially methylated regions in esophageal cancer reveal both cell-type- and cancer-specific epigenetic regulation 96%
Similar papers in this journal
- ModiDeC: a multi-RNA modification classifier for direct nanopore sequencing 97%
- A high-resolution map of functional miR-181 response elements in the thymus reveals the role of coding sequence targeting and an alternative seed match 96%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
Similar papers in this journal
Similar papers in this journal
- Multi-sample Full-length Transcriptome Analysis of 22 Breast Cancer Clinical Specimens with Long-Read Sequencing 96%
- Comprehensive evaluation and prediction of editing outcomes for near-PAMless adenine and cytosine base editors 95%
- Annotation of Chromatin States in 66 Complete Mouse Epigenomes During Development 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.