Discovery Stack Pilot: Feasibility and Outcomes of a Scientist-Designed Peer Review Model Separating Quality and Impact
McGargill, M. A.; Liu, B. C.; Kuhns, M. S.; Mucida, D.; Rauch, I.; Rodda, L. B.; Koch, M. A.; Gonzalez Velozo, H.; Cadwell, K.; Freedman, T. S.; Scharschmidt, T. C.; Sever, R.; Ordovas-Montanes, J.; Oberst, A.; Runnette, B.; Krummel, M. F.
10.1101/2025.10.31.685758 bioRxivShow abstract
Peer review serves as the cornerstone of scientific quality control. Yet, the current journal-centric system is hindered by long timelines, high publication costs, inconsistent review quality, systemic biases, and editorial gatekeeping. Notably, the system is built around misaligned measures of impact that are tethered to journal branding and conflate scientific rigor (Quality) with perceived significance (Impact). Here, we report findings from the Discovery Stack Pilot Study, which tested a scientist-designed, journal-independent peer review model. The Discovery Stack model integrates in-line reviewer comments to promote constructive, improvement-focused feedback and generates separate, multimodal assessments of scientific Quality and Impact. To examine its feasibility and effectiveness, manuscripts enrolled in the pilot were reviewed in parallel with traditional journal review. A total of 162 reviews were completed, and survey data from 86 participants were analyzed to evaluate the experience of both authors and reviewers. The results showed that reviewers effectively evaluated Quality and Impact as separate dimensions, with Quality scores being more consistent across reviewers than Impact scores. Importantly, participants strongly supported the core elements of the Discovery Stack model and expressed enthusiasm for its broader adoption to enhance transparency, efficiency, and value in peer review. Future studies will explore integrating this model into a digital platform for reviewing and curating scientific discoveries to improve the production and dissemination of high-quality research.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 92%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 88%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.