BLINK: Ultrafast tandem mass spectrometry cosine similarity scoring
Harwood, T. V.; Treen, D.; Wang, M.; de Jong, W.; Northen, T.; Bowen, B. P.
Show abstract
SummaryMetabolomics has a long history of using cosine similarity to match experimental tandem mass spectra to databases for compound identification. Here we introduce the Blur-and-Link (BLINK) approach for scoring cosine similarity. BLINK calculates substantially equivalent cosine similarity scores (>99% identification agreement) over 1000 times faster than commonly used loop-based implementations by bypassing fragment alignment and simultaneously scoring all pairs of spectra using sparse matrix operations. This performance improvement can enable calculations to be performed that would typically be limited by time and available computational resources. Availability and ImplementationBLINK is implemented in Python3 and is published under a modified open source license. Code and license are available on Github: https://github.com/biorack/blink
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extremely fast and accurate open modification spectral library searching of high-resolution mass spectra using feature hashing and graphics processing units 98%
- MSnbase, efficient and elegant R-based processing and visualisation of raw mass spectrometry data 96%
- Fast and memory efficient searching of large-scale mass spectrometry data using Tide 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 97%
- WiPP: Workflow for improved Peak Picking for Gas Chromatography-Mass Spectrometry (GC-MS) data 97%
- Robust Moiety Model Selection Using Mass Spectrometry Measured Isotopologues 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.