GeNePi: a GPU-enhanced Next Generation Bioinformatics Pipeline for Whole Genome Sequencing Analysis
Marangoni, S.; Furia, F.; Charrance, D.; Fant, A.; Di Dio, S.; Trova, S.; Spirito, G.; Musacchia, F.; Coppe, A.; Gustincich, S.; Cavalli, A.; Vecchi, M.; Landuzzi, F.
Show abstract
Next Generation Sequencing (NGS) has revolutionized genome biology, enabling the rapid sequencing of an entire human genome and facilitating the integration of Whole Genome Sequencing (WGS) into both research and clinical applications. The high-throughput nature of NGS and the complex data processing required has driven the need for advanced computational infrastructures to analyse these large datasets. The aim of this work is to introduce an innovative bioinformatic pipeline, named GeNePi, for the efficient and precise analysis of WGS short paired-end reads. Built on the Nextflow framework with a modular structure, GeNePi incorporates GPU-accelerated algorithms and supports multiple work-flow configurations. The pipeline automates the extraction of biologically relevant insights from raw WGS data, including: disease-related variants such as single nucleotide variants (SNVs), small insertions or deletions (INDELs), copy number variants (CNVs), and structural variants (SVs). Optimized for high-performance computing (HPC) environments, it takes advantage of job-scheduler submissions, parallelised processing, and tailored resource allocation for each analysis step. Tested on synthetic and real datasets, GeNePi accurately identifies genomic variants, with performances comparable to that of state-of-art tools. These features make GeNePi a valuable instrument for large-scale analyses in both research and clinical contexts, representing a key step towards the establishment of National Centers for Computational and Technological Medicine.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- unCOVERApp: an interactive graphical application for clinical assessment of sequence coverage at the base-pair level 97%
- MosaiCatcher v2: a single-cell structural variations detection and analysis reference framework based on Strand-seq 96%
- SVJedi: Genotyping structural variations with long reads 96%
Similar papers in this journal
- ILIAD: A suite of automated Snakemake workflows for processing genomic data for downstream applications 95%
- Rare Copy Number Variant analysis in case-control studies using SNP Array Data: a scalable and automated data analysis pipeline 95%
- CNVizard: a lightweight streamlit application for an interactive analysis of copy number variants 95%
Similar papers in this journal
- NGSTroubleFinder: A tool for detection and quantification of contamination and kinship across human NGS data 97%
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 95%
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
Similar papers in this journal
- A graph clustering algorithm for detection and genotyping of structural variants from long reads 98%
- CNVpytor: a tool for CNV/CNA detection and analysis from read depth and allele imbalance in whole genome sequencing 97%
- DivBrowse - interactive visualization and exploratory data analysis of variant call matrices 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.