Back

A Distributed Computing Solution for Privacy-Preserving Genome-Wide Association Studies

Brito, C.; Paulo, J.; Ferreira, P. G.

2024-01-16 bioinformatics
10.1101/2024.01.15.575678 bioRxiv
Show abstract

Breakthroughs in sequencing technologies led to an exponential growth of genomic data, providing unprecedented biological in-sights and new therapeutic applications. However, analyzing such large amounts of sensitive data raises key concerns regarding data privacy, specifically when the information is outsourced to third-party infrastructures for data storage and processing (e.g., cloud computing). Current solutions for data privacy protection resort to centralized designs or cryptographic primitives that impose considerable computational overheads, limiting their applicability to large-scale genomic analysis. We introduce GO_SCPLOWYOSAC_SCPLOW, a secure and privacy-preserving distributed genomic analysis solution. Unlike in previous work, GO_SCPLOWYOSAC_SCPLOW follows a distributed processing design that enables handling larger amounts of genomic data in a scalable and efficient fashion. Further, by leveraging trusted execution environments (TEEs), namely Intel SGX, GO_SCPLOWYOSAC_SCPLOW allows users to confidentially delegate their GWAS analysis to untrusted third-party infrastructures. To overcome the memory limitations of SGX, we implement a computation partitioning scheme within GO_SCPLOWYOSAC_SCPLOW. This scheme reduces the number of operations done inside the TEEs while safeguarding the users genomic data privacy. By integrating this security scheme in Glow, GO_SCPLOWYOSAC_SCPLOW provides a secure and distributed environment that facilitates diverse GWAS studies. The experimental evaluation validates the applicability and scalability of GO_SCPLOWYOSAC_SCPLOW, reinforcing its ability to provide enhanced security guarantees. Further, the results show that, by distributing GWASes computations, one can achieve a practical and usable privacy-preserving solution.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.