Back

MaveDB v2: a curated community database with over three million variant effects from multiplexed functional assays

Rubin, A. F.; Min, J. K.; Rollins, N. J.; Da, E. Y.; Esposito, D.; Harrington, M.; Stone, J.; Bianchi, A. H.; Dias, M.; Frazer, J.; Fu, Y.; Gallaher, M.; Li, I.; Moscatelli, O.; Ong, J. Y.; Rollins, J. E.; Wakefield, M. J.; Ye, S.; Tam, A.; McEwen, A. E.; Starita, L. M.; Bryant, V. L.; Marks, D. S.; Fowler, D. M.

2022-01-18 genomics
10.1101/2021.11.29.470445 bioRxiv
Show abstract

A central problem in genomics is understanding the effect of individual DNA variants. Multiplexed Assays of Variant Effect (MAVEs) can help address this challenge by measuring all possible single nucleotide variant effects in a gene or regulatory sequence simultaneously. Here we describe MaveDB v2, which has become the database of record for MAVEs. MaveDB now contains a large fraction of published studies, comprising over two hundred datasets and three million variant effect measurements. We created tools and APIs to streamline data submission and access, transforming MaveDB into a hub for the analysis and dissemination of these impactful datasets.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.