Reliable Molecular Retrieval from Mass Spectra using Conformal Prediction
Rakhshaninejad, M.; De Waele, G.; Jürgens, M.; Waegeman, W.
Show abstract
A key task in the computational analysis of liquid chromatography-tandem mass spectrometry (LC- MS/MS) data is identifying the molecular structure underlying a measured spectrum. A common approach ranks candidate molecules retrieved from chemical databases using predicted fingerprint similarities, yet standard metrics such as top-k accuracy summarize performance only at the dataset level and provide no spectrum-specific reliability statement. In this work, we apply conformal prediction to candidate-based molecular retrieval to construct spectrum-specific prediction sets that contain the true molecule with a user-specified probability. We evaluate marginal and conditional conformal prediction across three experimental scenarios representing in-distribution, partially shifted, and fully out-of-distribution settings on the MassSpecGym benchmark. When calibration and test data are aligned, conformal prediction attains the target coverage with small candidate sets for most spectra. Under distribution shift, prediction sets become larger as rankings grow more ambiguous, although candidates can still be reduced when calibration remains representative. Conditional conformal prediction improves subgroup reliability across spectra of different difficulty, with the best gains obtained using confidence-based grouping. Overall, conformal prediction turns candidate rankings into reliable, spectrum-specific candidate sets with an explicit reliability-efficiency trade-off.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- SingleFrag: A deep learning tool for MS/MS fragment and spectral prediction and metabolite annotation 94%
- ChemEmbed: A deep learning framework for metabolite identification using enhanced MS/MS data and multidimensional molecular embeddings 93%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.