Back

The Clinical Genomic Variation Landscape

Goar, W. A.; Pathawala, D.; Kuzma, K.; Bratulin, A.; Antoniou, A. A.; Arbesfeld, J. A.; Babb, L.; Ferriter, K.; O'Neill, T.; Stevenson, J. S.; Perry, K.; Canon, M.; Liu, J.; Liu, X.; Walsh, B.; Funk, S.; Ray, W. C.; Chaudhari, B. P.; Rehm, H. L.; Wagner, A. H.

2025-11-06 genetic and genomic medicine
10.1101/2025.11.04.25339115 medRxiv
Show abstract

Interpreting genomic variation requires analysts to collate and process information from disparate genomic evidence resources to discern the contributions to diseases and drug responses. Differences in variant representation across these evidence repositories includes nomenclature (e.g., HGVS, SPDI), reference sequence context (e.g., GRCh37 or GRCh38 genome assemblies), sequence annotation sources (e.g., RefSeq or Ensembl), and aggregate variant concepts (e.g., canonical alleles) collectively make it difficult to reveal whether (and how) genomic variants are associated with clinical outcomes. We evaluated these challenges across established genomic knowledge resources, including content from the CIViC, Molecular Oncology Almanac, and ClinVar knowledgebases, as compared against real-world small variant and CNV data. We used these findings to develop a suite of variant normalization methods to address these gaps. We present our findings as well as an analysis of remaining gaps in the representation of variation data and recommendations for the continued development of genomic knowledge standards to address these gaps.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.