Back

Mapping AAV capsid sequences to functions through function-guided in silico evolution

Zheng, H.; Guo, B.; Mo, A.; Wei, H.; Wu, Y.; Lin, X.; Jiang, H.; Li, H.; Zhang, Y.; Song, Z.; Ni, X.; Huang, Y.; Gu, X.; Yu, B.; Cheng, N.; Wang, X.

2024-10-11 bioengineering
10.1101/2024.10.11.617764 bioRxiv
Show abstract

Artificial intelligence (AI) offers significant potential to accelerate the functional engineering of adeno-associated virus (AAV) capsids in a time- and cost-efficient manner. However, existing approaches lack the ability to systematically map capsid sequences to multifunctional properties. To address this challenge, we developed ALICE, a knowledge-driven platform for in silico AAV capsid engineering that directly maps sequence to multiple functions. Our method incorporating a heuristic algorithm with contrastive learning principles, termed function-guided evolution (FE), to iteratively optimize high-performing capsid sequences generated by a naive language model toward desired functions. We elucidate FEs evolutionary mechanism, demonstrating its ability to navigate a function-guided landscape and generate capsids with tailored properties. After validating ALICEs ability to design murine-specific AAV capsids targeting Ly6a/Ly6c1 with limited training data, we developed ALICE-X by integrating a curriculum learning strategy. This upgrade facilitated exploration of sequence space beyond existing wet-lab datasets, enabling successful targeting of the human transferrin receptor 1 (hTfR1). The resulting engineered AAV variant, AAV.ALICE-H3, demonstrated improved viability, specific receptor targeting, and [~] 251-fold enhanced CNS tropism in human TFRC knock-in (hTFRC KI) mice compared to wildtype controls. This interpretable, knowledge-driven framework advances in silico AAV capsid design, offering broad applicability for translational gene therapy.

Published in Cell Press Blue · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.