Back

Topological stratification of continuous genetic variation in large biobanks

Diaz-Papkovich, A.; Zabad, S.; Ben-Eghan, C.; Anderson-Trocme, L.; Femerling, G.; Nathan, V.; Patel, J.; Gravel, S.

2023-07-07 genomics
10.1101/2023.07.06.548007 bioRxiv
Show abstract

Biobanks now contain genetic data from millions of individuals. Dimensionality reduction, visualization and clustering are standard when exploring data at these scales; while efficient and tractable methods exist for the first two, clustering remains challenging because of uncertainty about sources of population structure. In practice, clustering is commonly performed by drawing shapes around dimensionally reduced data or assuming populations have a "type" genome. We propose a method of clustering data with topological analysis that is fast, easy to implement, and integrates with existing pipelines. The approach is robust to the presence of sub-populations of varying sizes and wide ranges of population structure patterns. We use UMAP and HDBSCAN, respectively methods of dimensionality reduction and density clustering, on data from three biobanks. We illustrate how topological genetic strata can help us understand structure within biobanks, evaluate distributions of genotypic and phenotypic data, examine polygenic score transferability, identify potential influential alleles, and perform quality control.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.