SAKit: an all-in-one analysis pipeline for identifying novel protein caused by variant events at genomic and transcriptic level
Li, Y.; Wang, B.; Ding, W. Z.; Xu, S.; Cui, F.; Fei, C.; Sun, Q.
Show abstract
SummaryGenetic modifications that cause pivotal protein inactivation or abnormal activation may lead to cell signaling pathway change or even dysfunction, resulting in cancer and other diseases. In turn, dysfunction will further produce "novel proteins" that do not exist in the canonical human proteome. Identification of novel proteins is meaningful for identifying promising drug targets and developing new therapies. In recent years, several tools have been developed for identifying DNA or RNA variants with the extensive application of nucleotide sequencing technology. However, these tools mainly focus on point mutation and have limited performance in identifying large-scale variants as well as the integration of mutations. Here we developed a hybrid Sequencing Analysis bioinformatic pipeline by integrating all relevant detection Kits(SAKit): this pipeline fully integrates all variants at the genomic and transcriptomic level that may lead to the production of novel proteins defined as proteins with novel sequences compare to all reference sequences by comprehensively analyzing the long and short reads. The analysis results of SAKit demonstrate that large-scale mutations have more contribution to the production of novel proteins than point mutations, and long-read sequencing has more advantages in large-scale mutation detection. Availability and implementationSAKit is freely available on docker image (https://hub.docker.com/repository/docker/therarna/sakit), which is mainly implemented within a Snakemake framework in Python language.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 97%
- GSA: An Independent Development Algorithm for Calling Copy Number and Detecting Homologous Recombination Deficiency (HRD) from Target Capture Sequencing 95%
- APA-Scan: Detection and Visualization of 3'-UTR APA with RNA-seq and 3'-end-seq Data 95%
Similar papers in this journal
- Sashimi.py: a flexible toolkit for combinatorial analysis of genomic data 98%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 96%
- CHOmics: a web-based tool for multi-omics data analysis and interactive visualization in CHO cell lines 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.