ECCFP: a consecutive full pass based bioinformatic analysis for eccDNA identification using Nanopore sequencing data
Zhang, T.; Li, W.; Zeng, Q.; Zhang, J.; Miao, B.; Li, M.; Luo, J.; Liu, T.; Chen, S.; Wan, S.
Show abstract
It is commonly known that extrachromosomal circular DNA (eccDNA) has the potential as a molecular marker because of its close relationship with cancer progress and its prevalent existence in eukaryotic organisms. The mainstream technique of eccDNA detection is using high-throughput sequencing supported by bioinformatics analysis. Although these have various analysis pipelines for sequencing data, they are restricted by sequencing platforms or have shortcomings in accuracy and efficiency. To address these limitations, we design ECCFP, a bioinformatic analysis pipeline that detects eccDNAs amplified by rolling circle amplification (RCA) from long-read sequencing data and outputs eccDNA genomic coordinates and consensus sequences. This pipeline proposes a rigorous algorithm to retain all consecutive full passes derived from individual reads to obtain candidate eccDNAs, followed by systematic consolidation of candidate eccDNAs to detect unique eccDNAs. Using simulated datasets and experimental eccDNA sequencing datasets, we estimated ECCFP in several aspects and compared it with other existing pipelines. It exhibits a marked reduction in false positive rates compared with eccDNA_RCA_nanopore and superior sensitivity relative to CReSIL and FLED. Besides, inverse PCR and Sanger sequencing further validated the existence and accuracy of the position of detected eccDNAs by ECCFP. Collectively, ECCFP provides a more efficient choice for eccDNA detection from long-read sequencing data.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accurate Identification of Extrachromosomal Circular DNA from Long-read Sequences. 96%
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 96%
- LiBis: An ultrasensitive alignment method for low-input bisulfite sequencing 95%
Similar papers in this journal
- PtWAVE: A High-Sensitive deconvolution software of sequencing trace for the Detection of Large Indels in Genome Editing 97%
- DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing 96%
- Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP 96%
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 95%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
- Chromosome-scale assembly comparison of the Korean Reference Genome KOREF from PromethION and PacBio with Hi-C mapping information 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.