Back

Low-cost and highly efficient generation of near-complete bacterial pathogen genomes by TELL-Seq

Wu, Z.; Zhang, T.; Li, S.; Shen, L.; Yang, Z.; Yang, Y.; Chen, X.; Li, B.; Zhou, S.; Zhou, X.; Wu, B.; Jiang, J.; Li, X.

2024-08-19 bioinformatics
10.1101/2024.08.17.608388 bioRxiv
Show abstract

Recent evidences suggest that de novo genome assembly can provide additional insights into genomic variation landscape beyond short-read NGS analyses. Despite the advent of longread sequencing technologies, generating high-quality bacterial genome assemblies remains expensive and requires high-quality and large quantities of DNA. TELL-Seq, an emerging linked-read technology, has been reported as a low-cost alternative for producing nearcomplete genome assembly in certain bacterial species. However, a systematic assessment of its performance and characteristics in de novo bacterial genome assembly has not yet been conducted. To address this, we benchmarked TELL-Seq using a set of clinical and standard bacterial pathogens with a wide range of genome size (2.0 6.4 Mbp), GC-content (32% - 68%), genome complexity (mappability from 0.15 to 0.60) and Gram types. Our findings indicate that, with 1/6 cost and 1/15,000 of DNA compared to PacBio HiFi, TELL-Seq could generate as high quality near-complete bacterial genome assemblies. The optimal depth for assembly is around 200x, as increased depth improves contiguity and completeness, though quality plateaus at 250x. In general, genome complexity was significantly correlated with assembly quality, but GC-content was not. Nevertheless, genomes with extreme GC-content may require higher depths for accurate assembly. Our study suggests that TELL-Seq could be a scalable method for large-scale bacterial genomic surveys. The data generated from this study could serve as a benchmark dataset for further algorithm development. Impact StatementAn increasing body of literature have shown that de novo genome assembles could provide additional insights into genomic variation beyond standard short-read technologies. Its particularly useful in clinical and epidemiological studies of bacteria, due to high variable nature of their genomes. Although great advances have been made in long-read sequencing technology, its still expensive with high requirement over sample DNAs quality and quantity, limiting its adoption in large-scale studies or surveys. Linked-read technology is an inexpensive alternative which used single molecule barcoding to mark origin of reads for better scaffolding of contigs assembled from normal short reads. TELL-Seq, as one of these promising technologies, was reported to be able to inexpensively and efficiently generate near-complete bacterial de novo genomic assemblies with low DNA inputs. Theres not yet a systematic benchmarking of this technology in bacterial pathogens. In this study, a set of clinically important bacterial pathogens with a wide spectrum of GC-content, genome size, genome complexity, and Gram Type were used to investigate the performance of TELL-Seq technology in de novo genome assembly. We found that, with about 1/6 of cost and 1/15000 of DNA quantity, it could generate comparable de novo assemblies compared to standard reference genomes or PacBio HiFi assemblies, while beating conventional short-read assemblies. We found several key factors, such as genome complexity, and sequencing depth, which would impact the quality of the assembly. Our study suggests that a sequencing depth of 200x is sufficient to achieve satisfactory results. Our study has revealed characteristics of this technology in bacterial genome assembly and paved the way for its large-scale application in clinical or epidemiological surveys. Data SummaryO_LIRaw sequencing data for all isolates have been deposited to SRA, accessible through these NCBI BioProject PRJNA1133244. C_LIO_LIA full list of SRA BioSample accession numbers are available in Tab S5. C_LIO_LIAssemblies and analyses companion this paper is available at: https://github.com/x-lab/tellseq_bacteria C_LI

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.