Non-identifiability and the Blessings of Misspecification in Models of Molecular Fitness
Weinstein, E. N.; Amin, A. N.; Frazer, J.; Marks, D.
Show abstract
Understanding the consequences of mutation for molecular fitness and function is a fundamental problem in biology. Recently, generative probabilistic models have emerged as a powerful tool for estimating fitness from evolutionary sequence data, with accuracy sufficient to predict both laboratory measurements of function and disease risk in humans, and to design novel functional proteins. Existing techniques rest on an assumed relationship between density estimation and fitness estimation, a relationship that we interrogate in this article. We prove that fitness is not identifiable from observational sequence data alone, placing fundamental limits on our ability to disentangle fitness landscapes from phylogenetic history. We show on real datasets that perfect density estimation in the limit of infinite data would, with high confidence, result in poor fitness estimation; current models perform accurate fitness estimation because of, not despite, misspecification. Our results challenge the conventional wisdom that bigger models trained on bigger datasets will inevitably lead to better fitness estimation, and suggest novel estimation strategies going forward.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adaptive Estimation for Epidemic Renewal and Phylogenetic Skyline Models 97%
- ConvexML: Fast and accurate branch length estimation under irreversible mutation models, illustrated through applications to CRISPR/Cas9-based lineage tracing 96%
- Gene Transfer based Phylogenetics: Analytical Expressions and Additivity via Birth/Death Theory 96%
Similar papers in this journal
Similar papers in this journal
- Statistical inference of the rates of cell proliferation and phenotypic switching in cancer 96%
- From Bayes to Darwin: evolutionary search as an exaptation from sampling-based Bayesian inference 96%
- The ancestral population size conditioned on the reconstructed phylogenetic tree with occurrence data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.