Back

JUMP Cell Painting dataset: morphological impact of 136,000 chemical and genetic perturbations

Chandrasekaran, S. N.; Ackerman, J.; Alix, E.; Ando, D. M.; Arevalo, J.; Bennion, M.; Boisseau, N.; Borowa, A.; Boyd, J. D.; Brino, L.; Byrne, P. J.; Ceulemans, H.; Ch'ng, C.; Cimini, B. A.; Clevert, D.-A.; Deflaux, N.; Doench, J. G.; Dorval, T.; Doyonnas, R.; Dragone, V.; Engkvist, O.; Faloon, P. W.; Fritchman, B.; Fuchs, F.; Garg, S.; Gilbert, T. J.; Glazer, D.; Gnutt, D.; Goodale, A.; Grignard, J.; Guenther, J.; Han, Y.; Hanifehlou, Z.; Hariharan, S.; Hernandez, D.; Horman, S. R.; Hormel, G.; Huntley, M.; Icke, I.; Iida, M.; Jacob, C. B.; Jaensch, S.; Khetan, J.; Kost-Alimova, M.; Krawiec,

2023-03-24 bioinformatics
10.1101/2023.03.23.534023 bioRxiv
Show abstract

Image-based profiling has emerged as a powerful technology for various steps in basic biological and pharmaceutical discovery, but the community has lacked a large, public reference set of data from chemical and genetic perturbations. Here we present data generated by the Joint Undertaking for Morphological Profiling (JUMP)-Cell Painting Consortium, a collaboration between 10 pharmaceutical companies, six supporting technology companies, and two non-profit partners. When completed, the dataset will contain images and profiles from the Cell Painting assay for over 116,750 unique compounds, over-expression of 12,602 genes, and knockout of 7,975 genes using CRISPR-Cas9, all in human osteosarcoma cells (U2OS). The dataset is estimated to be 115 TB in size and capturing 1.6 billion cells and their single-cell profiles. File quality control and upload is underway and will be completed over the coming months at the Cell Painting Gallery: https://registry.opendata.aws/cellpainting-gallery. A portal to visualize a subset of the data is available at https://phenaid.ardigen.com/jumpcpexplorer/.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.