Back

Adversarial Attacks on Genotype Sequences

Mas Montserrat, D.; Ioannidis, A. G.

2022-11-08 bioinformatics
10.1101/2022.11.07.515527 bioRxiv
Show abstract

Adversarial attacks can drastically change the output of a method by performing a small change on its input. While they can be a useful framework to analyze worst-case robustness, they can also be used by malicious agents to perform damage in machine learning-based applications. The proliferation of platforms that allow users to share their DNA sequences and phenotype information to enable association studies has led to an increase in large databases. Such open platforms are, however, vulnerable to malicious users uploading corrupted genetic sequence files that could damage downstream studies. Such studies commonly include steps involving the analysis of the genomic sequences structure using dimensionality reduction techniques and ancestry inference methods. In this paper we show how white-box gradient-based adversarial attacks can be used to corrupt the output of genomic analyses, and we explore different machine learning techniques to detect such manipulations.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.