Predicted Effector Gene Aggregation, Standards and Unified Schema (PEGASUS): A Community Framework for Effector Gene Reporting
McMahon, A.; Ji, Y.; Costanzo, M.; Butterworth, A. S.; Pahl, M.; Szyszkowski, S.; Heilbron, K.; Shiyanbola, A.; Tsepilov, Y. A.; Spracklen, C. N.; Hite, D.; Shilin, A.; PEG Working Group, ; Parkinson, H. E.; Burtt, N. P.; Harris, L. W.
Show abstract
Genome-wide association studies (GWAS) increasingly report predicted effector genes (PEGs) - genes hypothesised to mediate the biological effects of associated variants. These function as key outputs for advancing variant-to-function research, mechanistic understanding, and therapeutic discovery. However, the rapid growth of PEG lists has not been matched by standards for organising, annotating, and reporting these predictions. As shown by recent landscape analyses, PEG lists vary widely in methodology, evidence definition, nomenclature, provenance tracking, and data structure, limiting interoperability, benchmarking, reuse, and adherence to FAIR principles. To address this gap, we convened an international multi-stakeholder community comprising method developers, data generators, resource maintainers, curators, funders, journal editors, and downstream users. Through a 2024 workshop and a 2025 working group series, we developed the Predicted Effector Gene Aggregation, Standards and Unified Schema (PEGASUS) framework. PEGASUS specifies (i) a metadata standard to capture provenance, trait and GWAS descriptors, evidence sources, and integration methods; (ii) a structured evidence matrix reporting all genes and all evidence underpinning prioritisation at each locus; and (iii) a concise PEG list that summarises author-prioritised genes linked transparently to underlying evidence. The framework balances transparency, machine readability, burden on submitters, and alignment with existing community standards. PEGASUS provides the first community-developed schema for reporting predicted effector genes and their supporting evidence. Adoption of this framework by authors will improve the comparability, reproducibility, and reusability of PEG outputs across studies, facilitating more robust biological inference, enabling cross-resource comparison of gene-prioritisation methods to support community benchmarking, and integration into downstream resources and analytical pipelines. PEGASUS-compliant data can be shared via the PEG Data Registry platform (https://kpndataregistry.org/peg), promoting re-use and establishing the basis for future integration with publicly shared GWAS data.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The NHGRI-EBI GWAS Catalog: standards for reusability, sustainability and diversity 95%
- Echtvar: Compressed variant representation for rapid annotation and filtering of SNPs and indels 93%
- FAVOR: Functional Annotation of Variants Online Resource and Annotator for Variation across the Human Genome 92%
Similar papers in this journal
Similar papers in this journal
- Predicting the functional impact of single nucleotide variants in Drosophila melanogaster with FlyCADD 92%
- Zebrafish Information Network, the knowledgebase for Danio rerio research 91%
- rvTWAS: identifying gene-trait association using sequences by utilizing transcriptome-directed feature selection 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.