Efficient Pangenome Construction through Alignment-Free Residue Pangenome Analysis (ARPA)
Lal, A.; Moustafa, A. M.; Planet, P.
Show abstract
Protein sequences can be transformed into vectors composed of counts for each amino acid (vector of Residue Counts; vRC) that are mathematically tractable and retain information about homology. We use vRCs to perform alignment-free, residue-based, pangenome analysis (ARPA; https://github.com/Arnavlal/ARPA). ARPA is 70-90 times faster at identifying homologous gene clusters compared to standard techniques, and offers rapid calculation, visualization, and novel phylogenetic approaches for pangenomes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 96%
- CONSULT: Accurate contamination removal using locality-sensitive hashing 95%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 94%
Similar papers in this journal
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 95%
- cognac: rapid generation of concatenated gene alignments for phylogenetic inferencefrom large whole genome sequencing datasets 95%
- PoMeLo: a systematic computational approach to predicting metabolic loss in pathogen genomes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.