Back

High-throughput protein target mapping enables accelerated bioactivity discovery for ToxCast and PFAS compounds

Yang, D.; Wang, X.; Liu, J.; Nair, P.; Sun, J.; Gong, Y.; Qian, X.; Cui, C.; Zeng, H.; Dong, A.; Harding, R. J.; Burgess-Brown, N.; Beyett, T. S.; Song, D.; Krause, H.; Diamond, M. L.; Bolhuis, D. L.; Brown, N. G.; Arrowsmith, C. H.; Edwards, A. M.; Halabelian, L.; Peng, H.

2025-03-25 pharmacology and toxicology
10.1101/2025.03.20.644436 bioRxiv
Show abstract

Chemical pollution is a global threat to human health, yet the toxicity mechanism of most contaminants remains unknown. Here, we applied an ultrahigh-throughput affinity-selection mass spectrometry (AS-MS) platform to systematically identify protein targets of prioritized chemical contaminants. After benchmarking the platform, we screened 50 human proteins against 481 prioritized chemicals, including 446 ToxCast chemicals and 35 per-and polyfluoroalkyl substances (PFAS). Among 24,050 interactions assessed, we discovered 35 novel interactions involving 14 proteins, with fatty acid-binding proteins (FABPs) emerging as the most ligandable protein family. Given this, we selected FABPs for further validation, which revealed a distinct PFAS binding pattern: legacy PFAS selectively bound to FABP1, whereas replacement compounds, PFECAs, unexpectedly interacted with all FABPs. X-ray crystallography further revealed that the ether group enhances molecular flexibility of alternative PFAS, to accommodate the binding pockets of FABPs. Our findings demonstrate that AS-MS is a robust platform for the discovery of novel protein targets beyond the scope of the ToxCast program and highlight the broader protein-binding spectrum of alternative PFAS as potential regrettable substitutes.

Published in Proceedings of the National Academy of Sciences (predicted rank #22) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.