InSilicoSeq 2.0: Simulating realistic amplicon-based sequence reads
Lelieveld, S. H.; Maas, T.; Duk, T. C. X.; Gourle, H.; van den Ham, H.-J.
Show abstract
MotivationSimulating high-throughput sequencing reads that mimic empirical sequence data is of major importance for designing and validating sequencing experiments, as well as for benchmarking bioinformatic workflows and tools. ResultsHere, we present InSilicoSeq 2.0, a software package that can simulate realistic Illumina-like sequencing reads for a variety of sequencing machines and assay types. InSilicoSeq now supports amplicon-based sequencing and comes with premade error models of various quality levels for Illumina MiSeq, HiSeq, NovaSeq and NextSeq platforms. It provides the flexibility to generate custom error models for any short-read sequencing platform from a BAM-file. We demonstrated the novel amplicon sequencing algorithm by simulating Adaptive Immune Receptor Repertoire (AIRR) reads. Our benchmark revealed that the simulated reads by InSilicoSeq 2.0 closely resemble the Phred-scores of actual Illumina MiSeq, HiSeq, NovaSeq and NextSeq sequencing data. InSilicoSeq 2.0 generated 15 million amplicon based paired-end reads in under an hour at a total cost of {euro}4.3e-05 per million bases advocating for testing experimental designs through simulations prior to actual sequencing. Availability and implementationInSilicoSeq 2.0 is implemented in Python and is freely available under the MIT licence at https://github.com/HadrienG/InSilicoSeq
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 93%
- ANOMALY: A Snakemake pipeline for identifying NuMTs from Long-Read Sequencing Data 92%
- Metagenomics-Toolkit: The Flexible and Efficient Cloud-Based Metagenomics Workflow featuring Machine Learning-Enabled Resource Allocation 92%
Similar papers in this journal
- Rescuing Low Frequency Variants within Intra-Host Viral Populations directly from Oxford Nanopore sequencing data 94%
- Comprehensive generation, visualization, and reporting of quality control metrics for single-cell RNA sequencing data 93%
- Automated strain separation in low-complexity metagenomes using long reads 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.