CanDIG: Secure Federated Genomic Queries and Analyses Across Jurisdictions
Dursi, L. J.; Bozoky, Z.; de Borja, R.; Li, J.; Bujold, D.; Lipski, A.; Rashid, S. F.; Sethi, A.; Memon, N.; Naidoo, D.; Coral-Sasso, F.; Wong, M.; Quirion, P.-O.; Lu, Z.; Agarwal, S.; Pavlov, K.; Ponomarev, A.; Husic, M.; Pace, K.; Palmer, S. L.; Grover, S. A.; Hakgor, S.; Siu, L. L.; Malkin, D.; Virtanen, C.; Pugh, T. J.; Jacques, P.-E.; Joly, Y.; Jones, S. J. M.; Bourque, G.; Brudno, M.
Show abstract
Rapid expansions of bioinformatics and computational biology have broadened the collection and use of -omics data including genomic, transcriptomic, methylomic and a myriad of other health data types, in the clinic and the laboratory. Both clinical and research uses of such data require co-analysis with large datasets, for which participant privacy and the need for data custodian controls must remain paramount. This is particularly challenging in multi-jurisdictional settings, such as Canada, where health privacy and security requirements are often heterogeneous. Data federation presents a solution to this, allowing for integration and analysis of large datasets from various sites while abiding by local policies. The Canadian Distributed Infrastructure for Genomics platform (CanDIG) enables federated querying and analysis of -omics and health data while keeping that data local and under local control. It builds upon existing infrastructures to connect five health and research institutions across Canada, relies heavily on standards and tooling brought together by the Global Alliance for Genomics and Health (GA4GH), implements a clear division of responsibilities among its participants and adheres to international data sharing standards. Participating researchers and clinicians can therefore contribute to and quickly access a critical mass of -omics data across a national network in a manner that takes into account the multi-jurisdictional nature of our privacy and security policies. Through this, CanDIG gives medical and research communities the tools needed to use and analyze the ever-growing amount of -omics data available to them in order to improve our understanding and treatment of various conditions and diseases. CanDIG is being used to make genomic and phenotypic data available for querying across Canada as part of data sharing for five leading pan-Canadian projects including the Terry Fox Comprehensive Cancer Care Centre Consortium Network (TF4CN) and Terry Fox PRecision Oncology For Young peopLE (PROFYLE), and making data from provincial projects such as POG (Personalized Onco- Genomics) more widely available.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sharing Data from the Human Tumor Atlas Network through Standards, Infrastructure, and Community Engagement 94%
- Nucleome Browser: An integrative and multimodal data navigation platform for 4D Nucleome 93%
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Orchestrating and sharing large multimodal data for transparent and reproducible research 95%
- A user's guide to the online resources for data exploration, visualization, and discovery for the Pan-Cancer Analysis of Whole Genomes project (PCAWG) 94%
- The 4D Nucleome Data Portal: a resource for searching and visualizing curated nucleomics data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.