Back

SpikeCleaner: An Algorithm to Label Unit Quality After Automated Spike Sorting

Zutshi, D.; Berezhnoi, D.; Ghimire, A.; Hartner, J.; Kim, D.; Watson, B. O.

2026-06-23 neuroscience
10.64898/2026.06.18.733033 bioRxiv
Show abstract

GapAutomated spike sorting algorithms have revolutionized the way neuronal activity is extracted from extracellular recordings, yet they remain imperfect. Specifically, inaccurate acceptance of noise-based units not only leaves researchers with clusters that require extensive manual curation, an essential but time-consuming process, that also leads to significant subjectivity in the selection of units. In an era of high-density probes like Neuropixels, where an hour of data can exceed 80 GB, manual curation is no longer scalable, automation of standard criteria can speed data curation and ensure quality of datasets. Here, we developed a semi-automated curation pipeline to label the quality of units after automated curation by Kilosort. ApproachOur algorithm standardizes criteria for labeling of Noise, Multi-Unit Activity (MUA), and Good Units using a combination of spike rate, spike timing metrics (from autocorrelogram), and waveform-based physiological features such as peak amplitude, slopes, half-width, and inter-channel correlation. Based on these features, clusters are assigned standardized labels (good, noise, multi-unit activity) that can be imported directly into Phy, where they serve as curation aids rather than absolute classifications, supporting but not replacing expert judgment. Heuristically, "noise" units are those unlikely to be neuronal in origin; "MUA" includes units with significant neural contribution (i.e., neuronal waveform) but with some degree of clear imperfection to be further cleaned, and "good" units are those without any clear deviation from ideal unit criteria. By ensuring accurate selection of acceptable units, we enable robust downstream analyses such as neural decoding and longitudinal tracking of neuron identity. Thresholds for all metrics were chosen to maximize the matching of algorithm output to that of 2 expert manual curators. Of note, users may alter thresholds either based on their own judgment or using an included tool to semi-automatically find thresholds that optimize SpikeCleaner with their own expert curation. Results: To benchmark, we compared the outputs of our algorithm to expert-labels curated in Phy by two expert users across three recordings. SpikeCleaner achieved an average of 97% accuracy vs. experts & 92% F1 score in classifying Single Units. It achieved an accuracy of 97% & 92% F1 score in full-category agreement (SU, MUA, Noise), and 97% accuracy & 95% F1 score in distinguishing Neuronal vs. Non-Neuronal units.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
eneuro
439 papers in training set
Top 0.1%
21.8%
2
eLife
5828 papers in training set
Top 10%
9.6%
3
Nature Neuroscience
252 papers in training set
Top 0.8%
7.8%
4
Wellcome Open Research
67 papers in training set
Top 0.1%
6.2%
5
PLOS Computational Biology
1863 papers in training set
Top 7%
5.4%
50% of probability mass above
6
Journal of Neural Engineering
221 papers in training set
Top 0.7%
4.8%
7
Neuron
337 papers in training set
Top 2%
4.8%
8
Journal of Neuroscience Methods
122 papers in training set
Top 0.5%
4.0%
9
Nature Communications
5641 papers in training set
Top 36%
3.2%
10
PLOS Biology
486 papers in training set
Top 2%
3.2%
11
Scientific Reports
3612 papers in training set
Top 43%
2.4%
12
PLOS ONE
5266 papers in training set
Top 42%
2.4%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 21%
2.4%
14
Cell Reports Methods
165 papers in training set
Top 3%
1.1%
15
Communications Biology
993 papers in training set
Top 22%
1.1%
16
Nature
645 papers in training set
Top 9%
1.1%
17
Imaging Neuroscience
282 papers in training set
Top 4%
1.0%
18
Scientific Data
209 papers in training set
Top 3%
0.8%
19
Bioinformatics
1204 papers in training set
Top 9%
0.8%
20
BMC Genomics
406 papers in training set
Top 8%
0.8%
21
Neuroinformatics
46 papers in training set
Top 1.0%
0.6%
22
Cell Reports
1498 papers in training set
Top 29%
0.6%
23
Nature Computational Science
55 papers in training set
Top 2%
0.6%
24
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.6%
25
npj Parkinson's Disease
105 papers in training set
Top 1%
0.6%
26
Nature Methods
385 papers in training set
Top 7%
0.6%
27
Journal of Visualized Experiments
34 papers in training set
Top 0.7%
0.6%
28
iScience
1154 papers in training set
Top 40%
0.6%