How to prepare the input data and run MCScanX efficiently?
Zhang, X.; Smith, D. R.
Show abstract
The protocol by Wang et al. is particularly useful as it outlines the steps for efficiently identifying colinear blocks in intra-and inter-species BLASTP outputs by using MCScanX. We recently discovered that the protocol lacks the pre-processing steps for checking if there are multiple isoforms derived from alternative splicing. Conserved sequences derived from alternative splicing can have similar functional domains, to avoid mis-prediction of gene duplicates, especially for the genome data from NCBI or other online resources. Without this step, the number of duplicate genes will be overrepresented. Besides we shared some useful experience to faster preparing the input data and easier running MCScanX. This is including alternative options to prepare the .gff input file and iterated all-against-all BLASTP processing. Lastly, we want to raise awareness of the potential challenges when preparing the input files and highlight potential issues when using the protocol.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NGScloud2: optimized bioinformatic analysis using Amazon Web Services 95%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 94%
- Non-synonymous to synonymous substitutions suggest that orthologs tend to keep their functions, while paralogs are a source of functional novelty 93%
Similar papers in this journal
- Dotplotic: a lightweight visualization tool for BLAST+ alignments and genomic annotations 95%
- ggcoverage: an R package to visualize and annotate genome coverage for various NGS data 94%
- Plant Co-expression Annotation Resource: a webserver for identifying targets for genetically modified crop breeding pipelines 93%
Similar papers in this journal
Similar papers in this journal
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 94%
- Comparison of visualisation tools for single-cell RNAseq data 93%
- Nextflow vs. plain Bash: Different Approaches to the Parallelisation of SNP Calling from the Whole Genome Sequence Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.