Back

Genome-Bench: A Scientific Reasoning Benchmark from Real-World Expert Discussions

Yin, M.; Qu, Y.; Liu, D.; Yang, L.; Cong, L.; Wang, M.

2025-06-05 genomics
10.1101/2025.06.02.657538 bioRxiv
Show abstract

In this short report, we present an automated pipeline tailored for the genomics domain and introduce Genome-Bench, a new benchmark constructed from over a decade of scientific forum discussions on genome engineering. Our pipeline transforms raw interactions into a reinforcement learningfriendly multiple-choice questions format, supported by 3000+ high-quality questionanswer pairs spanning foundational biology, experimental troubleshooting, tool usage, and beyond. To our knowledge, this is the first end-to-end pipeline for teaching LLMs to reason from scientific discussions, with promising potential for generalization across scientific domains beyond biology. The dataset is available at https://huggingface.co/datasets/Mingyin0312/Genome-Bench.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.