RL-Finetuning of OpenAI o1-mini to Enhance Biomedical Reasoning
Swanson, K.; Chen, Y. T.; Jaech, A.; Zou, J.
Show abstract
Recent breakthroughs in advanced reasoning large language models (LLMs), such as OpenAIs o1, have achieved impressive results in domains like math and coding. However, its not clear how much this type of reasoning helps in solving biomedical problems that involve more domain specialized knowledge and open-ended reasoning. Across two biomedical domains--gene characterization and small molecule property prediction--we find that the commercially available o1-mini model does not consistently outperform non-reasoning LLMs like GPT-4o. This motivated us to explore how much we can improve o1-minis biomedical reasoning through reinforcement learning (RL) finetuning. We show that RL finetuning of o1-mini results in large improvements in performance on gene classification, where it surprisingly outperformed domain-specific state-of-the-art models on some tasks. The results are mixed for small molecule prediction, suggesting that chemical reasoning could be more challenging for LLMs. We conclude with a discussion of the challenges and takeaways from this initial exploration of RL finetuning reasoning models for biomedical tasks.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 95%
- Accessible, Reproducible, and Scalable Machine Learning for Biomedicine 94%
- Causal reasoning over knowledge graphs leveraging drug-perturbed and disease-specific transcriptomic signatures for drug discovery 94%
Similar papers in this journal
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 94%
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 94%
- All-Atom Protein Sequence Design using Discrete Diffusion Models 94%
Similar papers in this journal
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 94%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.