Back

MutAIverse: An AI-Powered, Mechanism-backed Platform for Discovering Novel DNA Adducts and their precursor Genotoxins

Satija, S.; Mohanty, S. K.; Jorvekar, S.; Verma, A.; Solanki, S.; Gautam, V.; Arora, S.; Mittal, A.; Sharma, A.; Chauhan, S.; Duari, S.; Kumar, S.; Rai, A.; Das, A.; Kataki, K.; Das, K.; Sarma, A.; Tayal, J.; Sengupta, D.; Mehta, A.; Borkar, R. M.; Ahuja, G.

2025-08-30 bioinformatics
10.1101/2025.08.26.672508 bioRxiv
Show abstract

Genotoxin exposure leads to DNA adduct formation, potentially causing mutations if unrepaired. Current DNA adductomics platforms or analytical workflows are limited by incomplete spectral libraries, reliance on experimentally validated adducts, limited cellular contexts, and inefficient computational methodologies. We introduce MutAIverse, an advanced AI-driven DNA adductomics analysis platform that overcomes these limitations by leveraging intracellular mechanistic modeling of genotoxin bioactivation to construct a comprehensive DNA adduct library and offers advanced spectral mapping. MutAIverse integrates experimentally validated and chemically valid putative DNA adducts, enabling interpretable retrograde tracking to parental genotoxins. Validation of MutAIverse against experimental MS/MS datasets demonstrated the detection of both known and novel adducts. Furthermore, application to in-house generated DNA adductomics data from tissue biopsies of smokeless tobacco-induced head and neck cancer patients revealed selective enrichment of both novel and known DNA adducts; a subset of them was validated using MS/MS analysis. Collectively, MutAIverse provides a robust, end-to-end, and interpretable platform for advanced DNA adductomics analysis.

Published in Journal of Cheminformatics · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.