Back

CONCLAVE: CONsensus CLustering with Annotation-Validation Extrapolation for cyclic multiplexed immunofluorescence data

Nazari, P.; Arnould, A.; Andhari, M. D.; Fontecha, M.; Hernandez, J. M.; De Moor, B.; Pey, J.; De Smet, F.; Bosisio, F. M.; Antoranz, A.

2025-11-14 bioinformatics
10.1101/2025.11.13.688190 bioRxiv
Show abstract

High-dimensional cyclic multiplexed immunofluorescence (cMIF) enables single-cell phenotyping within intact tissues. Cell annotations rely on a multi-step pipeline involving normalization, sampling, dimensionality reduction, and clustering, but the absence of standardized benchmarks for method selection--especially at the clustering stage--leads to inconsistent and less reproducible phenotyping. To address this, we developed CONCLAVE, a consensus-clustering-based workflow that optimizes upstream steps and integrates results from multiple clustering algorithms retaining only those cell labels supported by at least two independent methods. Through in-silico simulations and real-world cMIF datasets, CONCLAVE consistently outperformed single-clustering-method approaches in accuracy, reproducibility, and robustness, with improvements becoming more evident when mapped within spatial tissue contexts. Additionally, CONCLAVE includes a scoring module that flags regions likely to contain unreliable or inconsistent data, facilitating targeted quality control. In summary, CONCLAVE offers a robust framework for cell annotation in cMIF datasets, enhancing the reliability of downstream spatial proteomics analyses. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=118 SRC="FIGDIR/small/688190v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@14ddfecorg.highwire.dtl.DTLVardef@1a83a17org.highwire.dtl.DTLVardef@17dfc74org.highwire.dtl.DTLVardef@493c07_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.