Back

The Genotype and Phenotypes in Families (GPF) platform manages the large and complex data at SFARI

Chorbadjiev, L.; Cokol, M.; Weinstein, Z.; Shi, K.; Fleisch, C.; Dimitrov, N.; Mladenov, S.; Xu, S.; Hall, J.; Ford, S.; Lee, Y.-h.; Yamrom, B.; Marks, S.; Munoz, A.; Lash, A.; Volfovsky, N.; Iossifov, I.

2024-02-11 bioinformatics
10.1101/2024.02.08.579330 bioRxiv
Show abstract

The exploration of genotypic variants impacting phenotypes is a cornerstone in genetics research. The emergence of vast collections containing deeply genotyped and phenotyped families has made it possible to pursue the search for variants associated with complex diseases. However, managing these large-scale datasets requires specialized computational tools tailored to organize and analyze the extensive data. GPF (Genotypes and Phenotypes in Families) is an open-source platform (https://github.com/iossifovlab/gpf) that manages genotypes and phenotypes derived from collections of families. The GPF interface allows interactive exploration of genetic variants, enrichment analysis for de novo mutations, and phenotype/genotype association tools. In addition, GPF allows researchers to share their data securely with the broader scientific community. GPF is used to disseminate two large-scale family collection datasets (SSC, SPARK) for the study of autism funded by the SFARI foundation. However, GPF is versatile and can manage genotypic data from other small or large family collections. Our GPF-SFARI GPF instance (https://gpf.sfari.org/) provides protected access to comprehensive genotypic and phenotypic data for the SSC and SPARK. In addition, GPF-SFARI provides public access to an extensive collection of de novo mutations identified in individuals with autism and related disorders and to gene-level statistics of the protected datasets characterizing the genes roles in autism. Here, we highlight the primary features of GPF within the context of GPF-SFARI.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.