Back

GWAC: A machine learning method to identify functional variants in data-constrained species

Sharo, A. G.

2024-11-17 genomics
10.1101/2024.11.15.623873 bioRxiv
Show abstract

As environments change, the ability of species to adapt depends on the functional variation they harbor. Identifying these functional variants is an important challenge in conservation genetics. Due to the limited data available for most species of conservation interest, genome-wide selection scans that link specific genetic variants with a phenotype are not feasible. However, functional variants may still be identified by considering predicted consequence, evolutionary conservation, and other sequence-based features. We developed Genome-Wide vAriant Classification (GWAC), a supervised machine learning framework to prioritize genome-wide variants by functional impact. GWAC requires only features that can be generated from an annotated genome. We evaluate GWAC by first using a set of human data constrained to match what may be available for threatened species. We find that GWAC weights features more heavily that are known to be predictive of functional variation and prioritizes both single nucleotide variants and indels, consistent with mutational constraint found in population genetics studies. GWAC performs nearly as well as CADD, a leading genome-wide predictor in humans that uses substantially more features and data that are typically available only for model organisms. While it is not possible to empirically evaluate GWAC on a species for which no functional variants are known, we find that a version of GWAC generated for the greater prairie chicken (Tympanuchus cupido pinnatus) weights features similarly to our human version. We compare the results of using a species-specific variant impact predictor against lifting-over variants from a closely related model organism and find that the species-specific approach retains functional variants that are lost during lift-over. We anticipate GWAC could be used to estimate conservation metrics such as genetic load and adaptive capacity, while also enabling researchers to identify individual variants responsible for adaptive phenotypes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.