Unlearning Virus Knowledge Toward Safe and Responsible Mutation Effect Predictions
Li, M.; Zhou, B.; Tan, Y.; Hong, L.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWPre-trained deep protein models have become essential tools in fields such as biomedical research, enzyme engineering, and therapeutics due to their ability to predict and optimize protein properties effectively. However, the diverse and broad training data used to enhance the generalizability of these models may also inadvertently introduce ethical risks and pose biosafety concerns, such as the enhancement of harmful viral properties like transmissibility or drug resistance. To address this issue, we introduce a novel approach using knowledge unlearning to selectively remove virus-related knowledge while retaining other useful capabilities. We propose a learning scheme, PROEDIT, for editing a pre-trained protein language model toward safe and responsible mutation effect prediction. Extensive validation on open benchmarks demonstrates that PROEDIT significantly reduces the models ability to enhance the properties of virus mutants without compromising its performance on non-virus proteins. As the first thorough exploration of safety issues in deep learning solutions for protein engineering, this study provides a foundational step toward ethical and responsible AI in biology.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 96%
- AITL: Adversarial Inductive Transfer Learning with input and output space adaptation for pharmacogenomics 96%
- TALE: Transformer-based protein function Annotation with joint sequence-Label Embedding 96%
Similar papers in this journal
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 95%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 95%
Similar papers in this journal
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 94%
- Evaluating generalizability of artificial intelligence models for molecular datasets 94%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.