Utanos: A general-purpose shallow whole-genome sequencing analysis workflow identifies interpretable copy number signatures
Douglas, J. M.; Lynch, B. J.; Yiu, J. C. H.; Nicholson, S.; Vasquez-Rios, C.; Ma, D.; Huntsman, D. G.; Park, Y.
Show abstract
SummaryA modular FASTQ-to-figures solution for analyzing low-depth or shallow whole-genome sequencing (sWGS) data. Shallow WGS can be used to detect copy number (CN) aberrations, Homologous Recombination Deficiency (HRD), and to detect and create CN signatures. Growing in popularity, this sequencing type is used for neonatal diagnostics and studying cancer. One of the major benefits is the reduced cost compared with deeper sequencing modalities such as Whole Genome Sequencing (WGS). With just 15 million reads often targeted, and sample substrate options like Formalin-Fixed, Paraffin-Embedded (FFPE) blocks widely available, this approach enables an affordable study to be performed at scale. Our pipeline and R package are an end-to-end solution implemented with reusability and modularity in mind. It makes entry and exit from the ecosystem easy, providing regular standardized output formats throughout execution. The pipeline is written in the well-supported and cross-platform Nextflow framework and has been submitted for inclusion in nf-core. Additionally, a Docker image for the utanos R package has been created to improve modularity. Availability and ImplementationThe latest version of all software is freely available on GitHub. For the full processing pipeline, visit: https://github.com/Huntsmanlab/swgs-processing-pipeline. For just the utanos R package, visit: https://github.com/Huntsmanlab/utanos.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 95%
- G2GSnake: A Snakemake workflow for host-pathogen genomic association studies 94%
- CLUES2 Companion: Computational pipelines to estimate, visualize, and date selection on multi-locus sites 94%
Similar papers in this journal
- ideal: an R/Bioconductor package for Interactive Differential Expression Analysis 93%
- cDNA-detector: Detection and removal of cDNA contamination in DNA sequencing libraries 93%
- CoGAPS 3: Bayesian non-negative matrix factorization for single-cell analysis with asynchronous updates and sparse data structures 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.