Distribution-Constrained Optimization for Reliable ML-Guided 5'UTR Sequence Design
Yamaguchi, R.; Mori, C.; Inoue, S.
Show abstract
The 5' untranslated region (5' UTR) shapes translation initiation, so its design is central to mRNA therapeutics and to improving protein-production cell lines. Deep-learning models that predict translation efficiency, measured as mean ribosome load (MRL), from the 5' UTR sequence have been combined with genetic algorithms (GAs) for sequence optimization. However, optimizing against a model trained on offline data risks reward hacking that exploits the models estimation error outside the training distribution, yielding sequences that score highly in prediction yet fail to perform in the wet lab. Yet for 5' UTR design, few studies have systematically examined which region should be treated as untrustworthy (the definition of out-of-distribution, OOD) or which constraints keep the search away from it. We present a constrained optimization that keeps candidates within a trust region where the predictors validated accuracy holds; here "reliable" denotes keeping candidates within the training distribution over which prediction has been validated, not a guarantee of measured performance. As the OOD score, we compare the k-nearest-neighbor (KNN) distance in the predictors embedding space against a pseudo-perplexity (PPPL) from the encoder and LM head, and show that for nucleotide sequences--whose vocabulary is small--PPPL fails to separate in- vs out-of-distribution, whereas the KNN distance is an effective OOD score that can define a trust region even from unlabeled native UTR sequences. Using the KNN distance as a hard GA constraint keeps all candidates inside the trust region while maintaining predicted MRL: under unconstrained optimization most final-generation candidates (72-96% across seeds) left the trust region (self-KNN p95), whereas the hard constraint holds predicted MRL at the unconstrained level and yields about 4.3x more selectable low-risk candidates than post-hoc filtering of the unconstrained output. Comparing an output extrapolation guard, reference-sequence similarity and structural accessibility (RNAplfold), we find that the guard and the similarity constraint also suppress OOD as a side effect, whereas making accessibility a secondary objective broadens the search without suppressing OOD.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Single-sequence protein-RNA complex structure prediction by geometric attention-enabled pairing of biological language models 93%
- Undersampling and the inference of coevolution in proteins 93%
- Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck 92%
Similar papers in this journal
- DUETT quantitatively identifies known and novel events in nascent RNA structural dynamics from chemical probing data 94%
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 94%
- Scalable Differentiable for mRNA Design 93%
Similar papers in this journal
Similar papers in this journal
- Target-site Dynamics and Alternative Polyadenylation Explain Large Share of Apparent MicroRNA Differential Expression 93%
- Deep Learning for RNA Secondary Structure Determination: Gauging Generalizability and Broadening the Scope of Traditional Methods 93%
- On the emergence of structural complexity in RNA replicators. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.