SpatialDataAgent: Autonomous Spatial Omics Data Curation at Decade Scale
Ji, J.-H.; Zou, Q.; Cheng, J.; She, Z.; Hao, Y.; Liu, W.; Zhang, D.; Wang, Z.; Yu, J.-T.; Yuan, Z.
Show abstract
Fragmented metadata in spatial omics archives has rendered large volumes of multimodal molecular-histological data inaccessible as dark data. Here, we introduce SpatialDataAgent, an agentic workflow for autonomous spatial omics data curation, combining schema-constrained evidence evaluation with a self-refining standardization agent. Applied to a decade of GEO records, SpatialDataAgent identified 769 paired H&E-spatial transcriptomics (ST) datasets, representing a 6.4-fold scale expansion over existing manually curated baselines. Within the benchmarking window, the framework achieved a 141% increase in high-confidence Class A paired datasets. We further assembled a recent high-confidence subset into HESRT, a standardized datalake containing 29.2 million spots/cells, establishing a blueprint for evidence-grounded autonomous curation of multimodal biomedical archives.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NiCo Identifies Extrinsic Drivers of Cell State Modulation by Niche Covariation Analysis 95%
- Characterizing cell-type spatial relationships across length scales in spatially resolved omics data 95%
- PHARAOH: A collaborative crowdsourcing platform for PHenotyping And Regional Analysis Of Histology 95%
Similar papers in this journal
Similar papers in this journal
- DeepSpaceDB: a spatial transcriptomics atlas for interactive in-depth analysis of tissues and tissue microenvironments 95%
- Probabilistic cell/domain-type assignment of spatial transcriptomics data with SpatialAnno 95%
- SPOTlight:Seeded NMF regression to Deconvolute Spatial Transcriptomics Spots with Single-Cell Transcriptomes 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.