Back

ortho_seqs: A Python tool for sequence analysis and higher order sequence-phenotype mapping

Nafees, S.; Vemuri, V. N.; Woollacott, M.; Solak, A. C.; Logan, P.; McGeever, A.; Yoo, O.; Rice, S. H.

2022-09-17 bioinformatics
10.1101/2022.09.14.506443 bioRxiv
Show abstract

MotivationAn important goal in sequence analysis is to understand how parts of DNA, RNA, or protein sequences interact with each other and to predict how these interactions result in given phenotypes. Mapping phenotypes onto underlying sequence space at first- and higher order levels in order to independently quantify the impact of given nucleotides or residues along a sequence is critical to understanding sequence-phenotype relationships. ResultsWe developed a Python software tool, ortho_seqs, that quantifies higher order sequence-phenotype interactions based on our previously published method of applying multivariate tensor-based orthogonal polynomials to biological sequences. Using this method, nucleotide or amino acid sequence information is converted to vectors, which are then used to build and compute the first- and higher order tensor-based orthogonal polynomials. We derived a more complete version of the mathematical method that includes projections that not only quantify effects of given nucleotides at a particular site, but also identify the effects of nucleotide substitutions. We show proof of concept of this method, provide a use case example as applied to synthetic antibody sequences, and demonstrate the application of ortho_seqs to other other sequence-phenotype datasets. Availabilityhttps://github.com/snafees/ortho_seqs & documentation https://ortho-seqs.readthedocs.io/

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.