Structured Chemical Reaction Modeling with Multitask Graph Neural Networks
Astero, M.; Li, A.; Casiraghi, E.; Rousu, J.
Show abstract
Modeling chemical reactions requires connecting fine-grained atom-bond edits with broader semantic categories. Yet, most machine learning approaches model these aspects in isolation: atom mapping, reaction center identification, and reaction classification are treated as separate problems. This separation limits accuracy, interpretability, and generalization. In this work, we argue that multitask learning provides a natural and powerful framework for reaction modeling. By jointly predicting mappings, centers, and classes within a single graph neural network, models can leverage structural dependencies between tasks. We proposed MARCC (Mapping-Assisted Reaction Center and Classification), a multitask architecture that achieves state-of-the-art performance on the USPTO-50K benchmark. MARCC demonstrates that multitask supervision not only improves accuracy across all tasks but also provides a structured representation of reaction mechanisms.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 95%
- All-Atom Protein Sequence Design using Discrete Diffusion Models 94%
- DrugDiff - small molecule diffusion model with flexible guidance towards molecular properties 94%
Similar papers in this journal
Similar papers in this journal
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 93%
- Benchmarking Uncertainty Quantification for Protein Engineering 93%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.