Context-dependent calibration of Evo2 likelihood with bacterial fitness: a quantitative characterization across five E. coli datasets
MinSeo, K.; Shin, J.-H.
Show abstract
DNA foundation models are trained to predict the likelihood of natural sequences, but the calibration between such likelihood scores and laboratory fitness or directly measured molecular phenotypes depends strongly on gene context, sequence divergence from wild-type, and selection regime. We apply zero-shot variant scoring with Evo2 7B (dLLR, the change in pseudo-log-likelihood between mutant and reference windows) to five E. coli datasets and quantify this context-dependent calibration map. Calibration is strong in two settings. In the Firnberg 2014 deep mutational scan of TEM-1 beta-lactamase (13,027 nucleotide-level variants; plasmid-borne enzyme under band-pass ampicillin selection), Evo2 dLLR tracks measured fitness at Spearman rho = 0.545 (95% CI 0.532-0.557; SNV rho = 0.606, indel rho = 0.521). In the Tenaillon 2012 thermal-evolution dataset, type-stratified, window-tuned scoring reaches Insertion AUROC 0.882 (W = 2,048 bp) and Deletion AUROC 0.846 (W = 4,096 bp). Calibration is decisively absent in the same organism: the Ireland 2020 RegSeq promoter MPRA gives rho = 0.011 (95% CI 0.003-0.019; n = 64,665), flat even after -10/-35 mechanism stratification, and the Dewachter 2023 chromosomal-essentials scan (fabZ/lpxC/murA) gives rho = 0.041 (95% CI 0.025-0.058). The Papkou 2023 folA combinatorial landscape sits between, at rho = 0.237, with a sweep that falls monotonically from rho = 0.575 at two mutations from wild-type to rho = 0.065 at nine. Pooling per-gene and per-divergence correlations, we fit calibration as an explicit function rho = f(sequence divergence from WT, variant context): weighted regression gives a negative divergence coefficient and a negative regulatory-context coefficient (both in the predicted direction; R2 = 0.49), an explicit, if illustrative, fit rather than a metaphor. We further test, and find unsupported, the intuitive explanation for the residual TEM-1 vs. essentials gap: across five genes the chromosomal essentials are more represented than TEM-1 by raw public-database deposition couno simple training over-representation does not exiant diversity is a candidate but remains untested. We therefore reframe Evo2 not a a likelihood predictor whosecalibration with fitness is conble is not a DMS pre-screen toolbut a quantitative lookup table likelihood-fitness gap closes(training-rich plasmid CDS undeens (chromosomal essentials,native promoter regulatory variorganism, plasmid vs. chromosomal context and strong vs. weak selifferent calibration regimes,the central finding.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Nucleotide-resolution bacterial pan-genomics with reference graphs 94%
- DiMSum: an error model and pipeline for analyzing deep mutational scanning data and diagnosing common experimental pathologies 94%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.