Back

Protein solubility depends on centrifugation: Aiki-Sol, a per-regime predictor for E. coli

Rajagopalan, R.; Meda, R. S.; Shastry, S.; Mysore, V.

2026-05-14 bioinformatics
10.64898/2026.05.14.725067 bioRxiv
Show abstract

MotivationSequence-based predictors of recombinant protein solubility in Escherichia coli have plateaued (NESG independent-test AUC 0.760 [->]~ 0.80 over eight years of protein-language-model variants). The plateau hides a latent confound: the centrifugation regime used to separate the soluble from the insoluble fraction is a hidden variable collapsed into a single binary "soluble" label. The proteins biochemistry does not change between regimes; what changes is which fraction of the lysate is recovered as soluble. Existing predictors treat the regime as label noise rather than a feature, and sequence overlap between training and test partitions masks the resulting failure mode. ResultsWe release the Aiki-Sol Dataset, a tiered E. coli solubility corpus: a ~ 85K stringency-annotated benchmark, an Apache-licensed ~ 147K extension adding binary-only-labelled proteins, and a ~ 229K research-tier pool incorporating non-commercially-licensed sources. On the ~ 85K benchmark, scored on sequence-cluster-disjoint partitions, the strongest published binary comparator falls below chance on the 32,000 x g stratum (AUC 0.491 {+/-} 0.020); a fine-tuned ESM-2 650M backbone with five protocol-matched out-puts lifts pooled AUC by +0.108 (paired-bootstrap CI lower bound +0.090). The gain is curation, not architecture: structure-aware predictors given ESMFold structures do not outperform the sequence-only frame, and capacity scaled to 3B parameters does not exceed the conditioned 650M backbone. The released model, Aiki-Sol, jointly supervises five per-stringency outputs alongside a marginal output for stringency-unknown proteins; on five external cohorts it lifts cohort-mean AUC from 0.69-0.70 to 0.825, with a [≥] +0.10-0.16 lift on the three cohorts at measurably-zero training-pool overlap. Availability and implementationAiki-Sol model weights (Apache 2.0), the 147K-row license-clean training pool of the deployment checkpoint (CC BY 4.0), the cluster-disjoint per-stringency 5-fold partition assignments, per-cohort prediction CSVs, and source code for training, inference, and figure reproduction are available at https://github.com/aikium-public/aiki-sol and archived at Zenodo 10.5281/zenodo.20151817. The research-tier 229K checkpoint is released under CC-BY-NC-ND 4.0 (inheriting the most-restrictive upstream-source tier); its training CSV and the 84,809-protein stringency-annotated bench-mark of [§]2.1 mix non-commercial-tier upstream sources and are not redistributed verbatim. Upstream sources are documented in Data availability and SI [§]S1. The deployment artefact is distributed as a Python package (pip install aikisol) with a predict(seq) entry point. Contactvenkatesh@aikium.com. Supplementary informationSupplementary text, figures, and tables are available at Bioinformatics online.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.