rbio1-training scientific reasoning LLMs with biological world models as soft verifiers
Istrate, A.-M.; Milletari, F.; Castrotorres, F.; Tomczak, J. M.; Torkar, M.; Li, D.; Karaletsos, T.
Show abstract
Reasoning models are typically trained against verification mechanisms in formally specified systems such as code or symbolic math. In open domains like biology, however, we lack exact rules to enable large-scale formal verification and instead often rely on lab experiments to test predictions. Such experiments are slow, costly, and cannot scale with computation. In this work, we show that world models of biology or other prior knowledge can serve as approximate oracles for soft verification, allowing reasoning systems to be trained without additional experimental data. We present two paradigms of training models with approximate verifiers: RLEMF: reinforcement learning with experimental model feedback and RLPK: reinforcement learning from prior knowledge. Using these paradigms, we introduce rbio1, a reasoning model for biology post-trained from a pretrained LLM with reinforcement learning, using learned biological models for verification during training. We demonstrate that soft verification can distill biological world models into rbio1, enabling it to achieve state-of-the-art performance on perturbation prediction in the PerturbQA benchmark. We further show that composing multiple AI-verifiers improves performance and that models trained with soft biological rewards transfer zero-shot to cross-domain tasks such as disease-state prediction. We present rbio1 as a proof of concept that predictions from biological models can train powerful reasoning systems using simulations rather than experimental data, offering a new paradigm for model training.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Learning Genetic Perturbation Effects with Variational Causal Inference 96%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 96%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 96%
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 96%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 96%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.