Back

Localized assembly for long reads enables genome-wide analysis of repetitive regions at single-base resolution in human genomes

Ikemoto, K.; Fujimoto, H.; Fujimoto, A.

2022-12-03 genomics
10.1101/2022.12.02.518938 bioRxiv
Show abstract

BackgroundLong-read sequencing technologies have the potential to overcome the limitations of short reads and provide a comprehensive picture of the human genome. However, it remains hard to characterize repetitive sequences by reconstructing genomic structures at high resolution solely from long reads. Here, we developed a localized assembly method (LoMA) that constructs highly accurate consensus sequences (CSs) from long reads. MethodsWe first developed LoMA, by combining minimap2, MAFFT, and our algorithm, which classifies diploid haplotypes based on structural variants and constructs CSs. Using this tool, we analyzed two human samples (NA18943 and NA19240) sequenced with the Oxford Nanopore sequencer. We defined target regions in each genome based on mapping patterns and then constructed a high-quality catalog of the human insertion solely from the long-read data. ResultsThe assessment of LoMA showed high accuracy of CSs (error rate < 0.3%) compared with raw data (error rate > 8%) and superiority to the previous study. The genome-wide analysis of NA18943 and NA19240 identified 5,516 and 6,542 insertions ({zeta} 100 bp) respectively. Most insertions ([~]80%) were derived from the tandem repeat and transposable elements. We also detected processed pseudogenes, insertions in transposable elements, and long insertions (> 10 kbp). Further, our analysis suggested that short tandem duplications were association with gene expression and transposons. ConclusionsOur analysis showed that LoMA constructs high-quality sequences from long reads with substantial errors. This study revealed the true structures of insertions with high accuracy and inferred mechanisms for the insertions. Our approach contributes to the future human genome studies. LoMA is available at our GitHub page: https://github.com/kolikem/loma.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Mobile DNA
31 papers in training set
Top 0.1%
26.7%
2
BMC Genomics
406 papers in training set
Top 0.3%
10.7%
3
Genome Biology
637 papers in training set
Top 1.0%
9.8%
4
Genomics, Proteomics & Bioinformatics
172 papers in training set
Top 0.2%
9.7%
50% of probability mass above
5
PLOS Genetics
862 papers in training set
Top 3%
4.3%
6
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.5%
7
Genome Research
468 papers in training set
Top 2%
3.2%
8
PLOS Computational Biology
1863 papers in training set
Top 11%
2.7%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.5%
10
Scientific Reports
3612 papers in training set
Top 43%
2.4%
11
Genes
144 papers in training set
Top 1%
2.4%
12
Nucleic Acids Research
1281 papers in training set
Top 8%
2.1%
13
GigaScience
212 papers in training set
Top 2%
1.7%
14
Genomics
64 papers in training set
Top 1.0%
1.4%
15
Frontiers in Genetics
230 papers in training set
Top 3%
1.4%
16
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
17
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
18
PLOS ONE
5266 papers in training set
Top 54%
1.1%
19
Nature Communications
5641 papers in training set
Top 53%
1.1%
20
eLife
5828 papers in training set
Top 62%
0.9%
21
Genome Medicine
183 papers in training set
Top 5%
0.8%
22
Bioinformatics
1204 papers in training set
Top 9%
0.8%
23
BMC Biology
265 papers in training set
Top 6%
0.6%
24
Journal of Genetics and Genomics
38 papers in training set
Top 0.9%
0.6%