Back

RL-Finetuning of OpenAI o1-mini to Enhance Biomedical Reasoning

Swanson, K.; Chen, Y. T.; Jaech, A.; Zou, J.

2025-05-24 bioinformatics
10.1101/2025.05.19.654988 bioRxiv
Show abstract

Recent breakthroughs in advanced reasoning large language models (LLMs), such as OpenAIs o1, have achieved impressive results in domains like math and coding. However, its not clear how much this type of reasoning helps in solving biomedical problems that involve more domain specialized knowledge and open-ended reasoning. Across two biomedical domains--gene characterization and small molecule property prediction--we find that the commercially available o1-mini model does not consistently outperform non-reasoning LLMs like GPT-4o. This motivated us to explore how much we can improve o1-minis biomedical reasoning through reinforcement learning (RL) finetuning. We show that RL finetuning of o1-mini results in large improvements in performance on gene classification, where it surprisingly outperformed domain-specific state-of-the-art models on some tasks. The results are mixed for small molecule prediction, suggesting that chemical reasoning could be more challenging for LLMs. We conclude with a discussion of the challenges and takeaways from this initial exploration of RL finetuning reasoning models for biomedical tasks.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.