GUANinE v1.1 Reveals Complementarity of Supervised and Genomic Language Models
robson, e. s.; Ioannidis, N. M.
Show abstract
BackgroundThere has been much debate about the benefits of supervised versus unsupervised learning on genomes. Answering the question of "which is better?" requires the development of comprehensive benchmarks spanning representative functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models under equivalent evaluation, such as L2-regularized probing (linear evaluation). ResultsHaving developed such an assessment (GUANinE v1.1), we can conclude that the answer is both: each paradigm outperforms on certain tasks and offers key advantages over the other. Supervised sequence-to-function models excel at annotating functional states characterized by chromatin accessibility, histone marks, or CTCF binding, while unsupervised language models outper-form on evolutionary conservation and related tasks without being limited to data-rich organisms. Our hundreds of new evaluations since v1.0 provide evidence for a direct tradeoff between input context size and model parameter count when on a fixed compute budget, which we depict with novel metrics like parameters/base pair. We also describe two new large-scale variant interpretation tasks: cadd-snv measuring proxy deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, language models, dominate deleteriousness prediction, but successfully translating deleteriousness to pathogenicity remains challenging. ConclusionsGUANinE v1.1 is a large-scale and thorough evaluation of pretrained genomic models. We identify uniquely performant models across tasks, and we conclude by suggesting hybrid-supervised language models may define the next era of genomic sequence modeling.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.