Back

A community driven GWAS summary statistics standard

Hayhurst, J.; Buniello, A.; Harris, L.; Mosaku, A.; Chang, C.; Gignoux, C. R.; Hatzikotoulas, K.; Karim, M. A.; Lambert, S. A.; Lyon, M.; McMahon, A.; Okada, Y.; Pirastu, N.; Rayner, N. W.; Schwartzentruber, J.; Vaughan, R.; Verma, S.; Wilder, S. P.; Cunningham, F.; Hindorff, L.; Wiley, K.; Parkinson, H.; Barroso, I.

2022-07-18 bioinformatics
10.1101/2022.07.15.500230 bioRxiv
Show abstract

Summary statistics from genome-wide association studies (GWAS) represent a huge potential for research. A challenge for researchers in this field is the access and sharing of summary statistics data due to a lack of standards for the data content and file format. For this reason, the GWAS Catalog hosted a series of meetings in 2021 with summary statistics stakeholders to guide the development of a standard format. The key requirements from the stakeholders were for a standard that contained key data elements to be able to support a wide range of data analyses, required low bioinformatics skills for file access and generation, to have easily accessible metadata, and unambiguous and interoperable data. Here, we define the specifications for the first version of the GWAS-SSF format, which was developed to meet the requirements discussed with the community. GWAS-SSF consists of a tab-separated data file with well-defined fields and an accompanying metadata file.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.