Back

SpatialDataAgent: Autonomous Spatial Omics Data Curation at Decade Scale

Ji, J.-H.; Zou, Q.; Cheng, J.; She, Z.; Hao, Y.; Liu, W.; Zhang, D.; Wang, Z.; Yu, J.-T.; Yuan, Z.

2026-05-30 bioinformatics
10.64898/2026.05.27.727615 bioRxiv
Show abstract

Fragmented metadata in spatial omics archives has rendered large volumes of multimodal molecular-histological data inaccessible as dark data. Here, we introduce SpatialDataAgent, an agentic workflow for autonomous spatial omics data curation, combining schema-constrained evidence evaluation with a self-refining standardization agent. Applied to a decade of GEO records, SpatialDataAgent identified 769 paired H&E-spatial transcriptomics (ST) datasets, representing a 6.4-fold scale expansion over existing manually curated baselines. Within the benchmarking window, the framework achieved a 141% increase in high-confidence Class A paired datasets. We further assembled a recent high-confidence subset into HESRT, a standardized datalake containing 29.2 million spots/cells, establishing a blueprint for evidence-grounded autonomous curation of multimodal biomedical archives.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.