Back

Metapipeline-DNA: A Comprehensive Germline & Somatic Genomics Nextflow Pipeline

Patel, Y.; Zhu, C.; Yamaguchi, T. N.; Wang, N.; Wiltsie, N.; Gonzalez, A.; Winata, H.; Zeltser, N.; Pan, Y.; Mootor, M. F. E.; Sanders, T.; Kandoth, C.; Fitz-Gibbon, S. T.; Livingstone, J.; Liu, L. Y.; Carlin, B.; Holmes, A.; Oh, J.; Sahrmann, J.; Tao, S.; Eng, S.; Hugh-White, R.; Pashminehazar, K.; Park, A.; Beshlikyan, A.; Jordan, M.; Wu, S.; Tian, M.; Arbet, J.; Neilsen, B.; Bugh, Y. Z.; Kim, G.; Salmingo, J.; Zhang, W.; Haas, R.; Anand, A.; Hwang, E.; Neiman-Golden, A.; Steinberg, P.; Zhao, W.; Anand, P.; Tsai, B. L.; Boutros, P. C.

2024-09-07 bioinformatics
10.1101/2024.09.04.611267 bioRxiv
Show abstract

SummaryThe price, quality and throughout of DNA sequencing continue to improve. Algorithmic innovations have allowed inference of a growing range of features from DNA sequencing data, quantifying nuclear, mitochondrial and evolutionary aspects of both germline and somatic genomes. To automate analyses of the full range of genomic characteristics, we created an extensible Nextflow meta-pipeline called metapipeline-DNA. Metapipeline-DNA analyzes targeted and whole-genome sequencing data from raw reads through pre-processing, feature detection by multiple algorithms, quality-control and data- visualization. Each step can be run independently and is supported robust software engineering including automated failure-recovery, robust testing and consistent verifications of inputs, outputs and parameters. Metapipeline-DNA is cloud-compatible and highly configurable, with options to subset and optimize each analysis. Metapipeline-DNA facilitates high-scale, comprehensive analysis of DNA sequencing data. AvailabilityMetapipeline-DNA is an open-source Nextflow pipeline under the GPLv2 license and is available at https://github.com/uclahs-cds/metapipeline-DNA.

Published in Cell Reports Methods (predicted rank #28) · training set

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.