DANCE 2.0: Transforming single-cell analysis from black box to transparent workflow
Ding, J.; Xing, Z.; Wang, Y.; Liu, R.; Liu, S.; Huang, Z.; Tang, W.; Xie, Y.; Zou, J.; Qiu, X.; Ma, J.; Yu, G.; Tang, J.
Show abstract
Preprocessing is a critical step in single-cell data analysis, yet current practices remain largely a black-box, trial-and-error process driven by user intuition, legacy defaults, and ad hoc heuristics. The optimal combination of steps such as normalization, gene selection, and dimensionality reduction varies across tasks, model architectures, and dataset characteristics, hindering reproducibility and method development. We present DANCE 2.0, an automated and interpretable preprocessing platform featuring two key modules: the Method-Aware Preprocessing (MAP) module, which discovers optimal pipelines for task-specific methods via hierarchical search, and the Dataset-Aware Preprocessing (DAP) module, which recommends pipelines for new datasets via similarity-based matching to a reference atlas. Together, MAP and DAP execute over 325,000 pipeline searches across six major tasks - clustering, cell type annotation, imputation, joint embedding, spatial domain identification, and cell type deconvolution - yielding robust and generalizable recommendations. MAP-recommended pipelines consistently outperform original method defaults, with substantial gains across all tasks. Beyond automation, DANCE 2.0 reveals interpretable preprocessing patterns across tasks, methods, and datasets, transforming preprocessing into a transparent, data-driven process. All resources are openly available at https://github.com/OmicsML/dance to support broad community adoption and future methodological advances.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.