Tox21mer, A transformer foundation model for Tox21 high-throughput concentration-response curves data
Li, L.; Hwang, J.; Shockley, K.; Li, Y.; Motsinger-Reif, A.; Hsieh, J.-H.; Auerbach, S. S.; Reif, D.
Show abstract
The U.S. Tox21 collaboration has generated a large reference library of high-throughput concentration-response assays. Here we present Tox21mer, a 43.5-million-parameter transformer that encodes each Tox21 concentration-response curve together with assay metadata into a 768-dimensional representation. Tox21mer was pretrained on ~2.5 million curves from 102 assay protocols and 6,727 compounds using masked-response reconstruction as the primary objective, with low-weight auxiliary supervision on assay outcome and AC50. To evaluate the learned representation, we trained lightweight probes on frozen embeddings from concentration-response curves of held-out compounds. The representation supported a macro-F1 of 0.985 for three-class outcome prediction (agonist, antagonist, inactive), a binary F1 of 0.994 for active/inactive prediction, and an R2 of 0.87 for log10(AC50). The learned embeddings formed coherent groupings by curve-class category. A masked-only pretraining variant retained near-baseline probe performance, indicating that the representation is learned largely from the self-supervised objective rather than from auxiliary labels. Ablation analyses further showed that predictive performance depends mainly on curve-level response-value distributions conditioned on assay context, with limited reliance on detailed within-curve ordering. Tox21mer thus provides a reusable foundation representation for Tox21 concentration-response data that can support extrapolation to untested compounds through integration with chemical features or distillation into chemistry-only student models for large-scale external screening.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 94%
- Trans-channel fluorescence learning improves high-content screening for Alzheimer's disease therapeutics 93%
- Sagittarius: Extrapolating Heterogeneous Time-Series Gene Expression Data 93%
Similar papers in this journal
Similar papers in this journal
- Adding stochastic negative examples into machine learning improves molecular bioactivity prediction 97%
- Compound activity prediction with dose-dependent transcriptomic profiles and deep learning 96%
- BOLD-GPCRs: A Transformer-Powered App for Predicting Ligand Bioactivity and Mutational Effects Across Class A GPCRs 95%
Similar papers in this journal
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 94%
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 93%
- A curriculum learning approach to training antibody language models 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.