Back

Efficient Pangenome Construction through Alignment-Free Residue Pangenome Analysis (ARPA)

Lal, A.; Moustafa, A. M.; Planet, P.

2022-06-05 bioinformatics
10.1101/2022.06.03.494761 bioRxiv
Show abstract

Protein sequences can be transformed into vectors composed of counts for each amino acid (vector of Residue Counts; vRC) that are mathematically tractable and retain information about homology. We use vRCs to perform alignment-free, residue-based, pangenome analysis (ARPA; https://github.com/Arnavlal/ARPA). ARPA is 70-90 times faster at identifying homologous gene clusters compared to standard techniques, and offers rapid calculation, visualization, and novel phylogenetic approaches for pangenomes.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.