Benchmarking DNA Foundation Models: Biological Blind Spots inEvo2 Variant-Effect Prediction
Mathur, V.; Sachidanandam, R.
Show abstract
DNA foundation models such as Evo and DNABERT-2 have generated considerable interest for clinically relevant genomics applications, particularly variant-effect prediction (VEP). However, rigorous benchmarks tailored to assessing their understanding of known biological constraints remain limited. Here, we develop controlled evaluation metrics based on well-characterized nuclear and mitochondrial sequence features and curated variant sets. Applying these benchmarks to Evo2, we identify systematic blind spots in short-range biological signals (e.g. codon usage bias) and observe sensitivity to contextual features that should be biologically neutral. These findings challenge current claims of zero-shot pathogenicity prediction and raise concerns regarding the clinical readiness of such models. The bench-marking framework introduced here provides a foundation for improving training strategies and for standardized evaluation of future genomic foundation models.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 94%
- Variant effect predictor correlation with functional assays is reflective of clinical classification performance 94%
- DiMSum: an error model and pipeline for analyzing deep mutational scanning data and diagnosing common experimental pathologies 94%
Similar papers in this journal
Similar papers in this journal
- Investigating the performance of foundation models on human 3'UTR sequences 93%
- tRNAscan-SE 2.0: Improved Detection and Functional Classification of Transfer RNA Genes 93%
- A mutational gradient drives somatic mutation accumulation in mitochondrial DNA and influences germline polymorphisms and genome composition 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.