Back

CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop

Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.

2026-08-20 bioinformatics
10.64898/2026.08.15.745007 bioRxiv
Show abstract

Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.