An End-to-end Oxford Nanopore Basecaller Using Convolution-augmented Transformer
Lv, X.; Chen, Z.; Lu, Y.; Yang, Y.
Show abstract
Oxford Nanopore sequencing is fastly becoming an active field in genomics, and its critical to basecall nucleotide sequences from the complex electrical signals. Many efforts have been devoted to developing new basecalling tools over the years. However, the basecalled reads still suffer from a high error rate and slow speed. Here, we developed an open-source basecalling method, CATCaller, by simultaneously capturing global context through Attention and modeling local dependencies through dynamic convolution. The method was shown to consistently outper-form the ONT default basecaller Albacore, Guppy, and a recently developed attention-based method SACall in read accuracy. More importantly, our method is fast through a heterogeneously computational model to integrate both CPUs and GPUs. When compared to SACall, the method is nearly 4 times faster on a single GPU, and is highly scalable in parallelization with a further speedup of 3.3 on a four-GPU node.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepGene: An Efficient Foundation Model for Genomics based on Pan-genome Graph Transformer 95%
- Trans-Driver: a deep learning approach for cancer driver gene discovery with multi-omics data 94%
- iDRKAN: Interpretable miRNA-Disease Association Prediction Based on Dual-Graph Representation Learning and Kolmogorov-Arnold Network 94%
Similar papers in this journal
Similar papers in this journal
- BertNDA: a Model Based on Graph-Bert and Multi-scale Information Fusion for ncRNA-disease Association Prediction 94%
- scIDPMs: Single-cell RNA-seq imputation using diffusion probabilistic models 94%
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.