Back

Community-Driven Copy Number Variant Discovery at Scale: Results from a Rare Disease Genomics Hackathon

Lun, M. Y.; Posey, J. E.; Bengtsson, J. D.; Du, H.; Roy, R. S.; Yang, L.; Ochoa, S.; Yuan, B.; Gillentine, M.; Lindstrand, A.; Carvalho, C. M. B.

2025-08-12 genetic and genomic medicine
10.1101/2025.08.08.25333317 medRxiv
Show abstract

PurposeCopy number variants (CNVs) are a major contributor to rare genetic diseases, but their detection and interpretation from short-read genome sequencing (srGS) data remain challenging, especially at scale. Large amounts of existing srGS data remain under-analyzed for clinically relevant CNVs. MethodsDuring a collaborative Hackathon, we developed and applied scalable CNV analysis workflows to srGS data from three unsolved, exome-negative, rare disease cohorts: Primary Immunodeficiency (N = 39), Turkish developmental disorders (N = 31), and data from the Genomics Research to Elucidate the Genetics of Rare diseases (GREGoR) (N = 1437). We employed Parliament2 for structural variant (SV) calling, Mosdepth and SLMSuite for read-depth-based quality control and CNV detection, and R Shiny-based visualization tools. We also constructed an SV/CNV variant database with population frequency and pathogenicity annotations, applied DBSCAN clustering for internal allele frequency estimation, and used a 3-way annotation strategy to aid interpretation. ResultsOur pipelines identified high-confidence CNVs and streamlined interpretation across cohorts. Within 2 days, the Hackathon yielded 39 candidate pathogenic SVs. The tools and workflows enabled rapid filtering, prioritization, and visualization of clinically relevant variants. ConclusionThis community-driven effort demonstrates the feasibility and utility of scalable CNV analysis for accelerating diagnosis and discovery in rare disease cohorts using srGS data.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.