Back

Structural variation across 138,134 samples in the TOPMed consortium

Jun, G.; English, A. C.; Metcalf, G. A.; Yang, J.; Chaisson, M. J.; Pankratz, N.; Menon, V. K.; Salerno, W. J.; Krasheninina, O.; Smith, A. V.; Lane, J. A.; Blackwell, T.; Kang, H. M.; Salvi, S.; Meng, Q.; Shen, H.; Pasham, D.; Bhamidipati, S.; Kottapalli, K.; Arnett, D. K.; Ashley-Koch, A.; Auer, P. L.; Beutel, K. M.; Bis, J. C.; Blangero, J.; Bowden, D. W.; Brody, J. A.; Cade, B. E.; Chen, Y.-D. I.; Cho, M. H.; Curran, J. E.; Fornage, M.; Freedman, B. I.; Fingerlin, T.; Gelb, B. D.; Hou, L.; Hung, Y.-J.; Kane, J. P.; Kaplan, R.; Kim, W.; Loos, R. J. F.; Marcus,, G. M.; Mathias, R. A.; McGarv

2023-01-26 genomics
10.1101/2023.01.25.525428 bioRxiv
Show abstract

Ever larger Structural Variant (SV) catalogs highlighting the diversity within and between populations help researchers better understand the links between SVs and disease. The identification of SVs from DNA sequence data is non-trivial and requires a balance between comprehensiveness and precision. Here we present a catalog of 355,667 SVs (59.34% novel) across autosomes and the X chromosome (50bp+) from 138,134 individuals in the diverse TOPMed consortium. We describe our methodologies for SV inference resulting in high variant quality and >90% allele concordance compared to long-read de-novo assemblies of well-characterized control samples. We demonstrate utility through significant associations between SVs and important various cardio-metabolic and hemotologic traits. We have identified 690 SV hotspots and deserts and those that potentially impact the regulation of medically relevant genes. This catalog characterizes SVs across multiple populations and will serve as a valuable tool to understand the impact of SV on disease development and progression.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.