Identifying parsimonious pathways of accumulation and convergent evolution from binary data
Giannakis, K.; Aga, O. N. L.; Moen, M. T.; Drange, P. G.; Johnston, I.
Show abstract
How stereotypical, and hence predictable, are evolutionary and accumulation dynamics? Here we consider processes - from genome evolution to cancer progression - involving the irreversible accumulation of binary features (characters), which can be modelled as Markov processes on a hypercubic transition network. We seek subgraphs of such networks that can generate a given set of paired before-after observations and minimize a topological cost function, involving criteria on out-branching which are interpretable in terms of biological parsimony. A transition network supporting a single, deterministic dynamic pathway is maximally simple and lowest cost, and branches (corresponding to possibly different next steps) increase cost, particularly if these branches are "deep", occurring at early stages in the dynamics. In this sense, the lowest-cost subgraph measures how stereotypical the evolutionary or accumulation process is, and also identifies good start points for likelihood-based inference. The problem is solvable in polynomial time for cross-sectional observations by building on an existing method due to Gutin, and we provide a polynomial-time estimate in the more general case of pairs of observed states. We use this approach to define a "stereotypy index" reflecting the extent of evolutionary predictability. We demonstrate use cases in the evolution of antimicrobial resistance, organelle genomes, and cancer progression, and provide a software implementation at https://github.com/StochasticBiology/hyperDAGs.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Squirrel: Reconstructing semi-directed phylogenetic level-1 networks from four-leaved networks or sequence alignments 96%
- Bayesian phylodynamic inference of multi-type population trajectories using genomic data 95%
- Weighting by Gene Tree Uncertainty Improves Accuracy of Quartet-based Species Trees 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.