Back

A robust unsupervised clustering approach for high-dimensional biological imaging data reveals shared drug-induced morphological signatures

Bao, S. C.; Mizikovsky, D.; Pishas, K.; Zhao, Q.; Cowley, K. J.; Marinovic, E.; Carey, M.; Campbell, I.; Simpson, K. J.; Cheasley, D.; Palpant, N.

2024-09-09 bioinformatics
10.1101/2024.09.05.611300 bioRxiv
Show abstract

Modern biology increasingly relies on large-scale screening to generate high dimensional datasets with potential to accelerate discovery. However, analysing these complex datasets remains challenging, particularly in applications where the underlying structure and groupings are unknown, and high dimensionality introduces noise and artifacts that make follow up studies difficult to prioritise. Here, we present an unsupervised consensus clustering tool that quantifies biologically meaningful patterns based on multi-scale data organisation to guide decision-making in high-throughput screening. Using large-scale drug screening data in cancer cell lines and bacterium model, we demonstrate its ability to use diverse data inputs to prioritize robust drug clusters with shared biological mechanisms and conserved drug responses. This method addresses key limitations associated with prioritising robust, actionable hits from scalable screening data.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.