Back

A simple way to find related sequences with position-specific probabilities

Frith, M.

2025-03-17 bioinformatics
10.1101/2025.03.14.643233 bioRxiv
Show abstract

One way to understand biology is by finding genetic sequences that are related to each other. Often, a family of related sequences has position-varying probabilities of substitutions, insertions, and deletions: we can use these to find distantly-related sequences. There are popular software tools for this task, which all have limitations. They either do not use all probability evidence (e.g. PSI-BLAST, MMseqs2), or have excessive complexity and minor biases (e.g. HMMER). This complexity inhibits fertile development of alternative tools. This study describes a simplest reasonable way to find related sequences, making full use of position-varying probabilities. The algorithms likely use the fewest operations that such algorithms possibly could, so they are fast and simple. This has been implemented in prototype software named DUMMER (Dumb Uncomplicated Match ModelER). Its sensitivity and specificity are competitive with HMMER. It finds evidence that the human genome has vastly more relics of LF-SINE, retrotransposons that were co-opted for various functions in common ancestors of all land vertebrates.

Published in Journal of Computational Biology (predicted rank #9) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.