Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics
Benegas, G.; Eraslan, G.; Song, Y. S.
Show abstract
Machine learning holds immense promise in biology, particularly for the challenging task of identifying causal variants for Mendelian and complex traits. Two primary approaches have emerged for this task: supervised sequence-to-function models trained on functional genomics experimental data and self-supervised DNA language models that learn evolutionary constraints on sequences. However, the field currently lacks consistently curated datasets with accurate labels, especially for non-coding variants, that are necessary to comprehensively benchmark these models and advance the field. In this work, we present TraitGym, a curated dataset of regulatory genetic variants that are either known to be causal or are strong candidates across 113 Mendelian and 83 complex traits, along with carefully constructed control variants. We frame the causal variant prediction task as a binary classification problem and benchmark various models, including functional-genomics-supervised models, self-supervised models, models that combine machine learning predictions with curated annotation features, and ensembles of these. Our results provide insights into the capabilities and limitations of different approaches for predicting the functional consequences of non-coding genetic variants. We find that alignment-based models CADD and GPN-MSA compare favorably for Mendelian traits and complex disease traits, while functional-genomics-supervised models Enformer and Borzoi perform better for complex non-disease traits. Evo2 shows substantial performance gains with scale, but still lags somewhat behind alignment-based models, struggling particularly with enhancer variants. The benchmark, including a Google Colab notebook to evaluate a model in a few minutes, is available at https://huggingface.co/datasets/songlab/TraitGym.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generative Haplotype Prediction Outperforms Statistical Methods for Small Variant Detection in NGS Data 96%
- RapidoPGS: A rapid polygenic score calculator for summary GWAS data without a test dataset 95%
- AdaLiftOver: High-resolution identification of orthologous regulatory elements with adaptive liftOver 95%
Similar papers in this journal
Similar papers in this journal
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 96%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 94%
- A global high-density chromatin interaction network reveals functional long-range and trans-chromosomal relationships 94%
Similar papers in this journal
- Best practices for multi-ancestry, meta-analytic transcriptome-wide association studies: lessons from the Global Biobank Meta-analysis Initiative 95%
- A Unifying Statistical Framework to Discover Disease Genes from GWAS 95%
- Fast and powerful statistical method for context-specific QTL mapping in multi-context genomic studies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.