MolSpecFlow: Mass-Constrained Hybrid Flow Matching for Joint Molecular-Spectral Analysis
Wang, Y.; Yang, F.; Xu, K.; Yuan, L.; Zhu, J.; Zhang, J.; Tang, Z.; Bian, Y.; Chang, C.; Tian, Y.; Yao, J.
Show abstract
Identifying the "dark matter" of the chemical universe requires bridging the fundamental gap between molecules and mass spectra. Existing approaches struggle to accurately map between these distinct data formats, often yielding results that are either chemically implausible or inconsistent with the physical evidence. We introduce MolSpecFlow, a unified foundation model pretrained on 100 million molecules and 42 million spectra, which leverages a hybrid flow matching framework to orchestrate optimal transport paths for spectral peaks and discrete probability flows for molecular tokens. Furthermore, to guarantee the physicochemical validity of the generated structures, we incorporate explicit rule-based constraints into the generative process, utilizing a token-level mass control mechanism that strictly enforces alignment with the precursor molecular weight. MolSpecFlow establishes a new state-of-the-art on the MassSpecGym benchmark, surpassing diffusion and retrieval baselines in de novo generation (Top-1 accuracy +35%), spectral simulation, and molecular retrieval, demonstrating the power of physics-grounded unified modeling.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data 95%
- Annotating metabolite mass spectra with domain-inspired chemical formula transformers 94%
- Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks 94%
Similar papers in this journal
Similar papers in this journal
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 94%
- Small molecule machine learning: All models are wrong, some may not even be useful 94%
- MolDiscovery: Learning Mass Spectrometry Fragmentation of Small Molecules 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.