Back

zsasa: a Zig-based engine for high-throughput solvent accessible surface area at proteome scale

Nagae, T.; Tomii, K.

2026-07-03 bioinformatics
10.64898/2026.06.29.733683 bioRxiv
Show abstract

Solvent accessible surface area (SASA) is widely used to describe protein stability, ligand binding, mutation effects, and protein-protein interfaces. As structural biology workloads expand to predicted-structure collections, trajectories, and large assemblies, SASA tools must combine reproducible calculation with high throughput, low memory use, and workflow-friendly input handling. We present zsasa, a Zig-based SASA engine with command-line and Python interfaces. zsasa implements the established Shrake-Rupley and Lee-Richards algorithms, provides exact f64/f32 modes and an optional bitmask approximation, and supports batch and trajectory workflows, compressed structure inputs, and configurable atom classification including Chemical Component Dictionary (CCD)-based radii for non-standard components. In matched Shrake-Rupley validation on 4,370 Escherichia coli AlphaFold Database structures, exact double-precision zsasa reproduced FreeSASA total SASA values to near numerical identity. In 10-thread batch benchmarks on the E. coli and 23,586-structure human AlphaFold collections, zsasa was 2.94x faster than a FreeSASA batch wrapper in exact f64 mode and up to 9.70x faster in bitmask mode, with roughly 4-8x lower peak memory. Trajectory benchmarks exceeded 1,000 frames/s at tens of megabytes of peak memory, and a 4.5-million-atom PDB stress-test file completed in under five seconds. These results support zsasa as a practical tool for reproducible, low-memory generation of surface-derived structural features at large scale. zsasa is available under the MIT License at https://github.com/N283T/zsasa.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Methods
385 papers in training set
Top 0.3%
19.5%
2
Bioinformatics
1204 papers in training set
Top 1%
19.5%
3
Protein Science
246 papers in training set
Top 0.1%
13.4%
50% of probability mass above
4
Journal of Molecular Biology
232 papers in training set
Top 0.4%
5.1%
5
Nature Communications
5641 papers in training set
Top 30%
4.6%
6
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.4%
7
Bioinformatics Advances
203 papers in training set
Top 2%
3.4%
8
Structure
193 papers in training set
Top 1%
2.2%
9
Nature Biotechnology
172 papers in training set
Top 2%
2.2%
10
Nucleic Acids Research
1281 papers in training set
Top 8%
2.0%
11
PLOS Computational Biology
1863 papers in training set
Top 14%
1.8%
12
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
1.8%
13
Cell Systems
201 papers in training set
Top 3%
1.6%
14
PLOS ONE
5266 papers in training set
Top 52%
1.4%
15
Communications Biology
993 papers in training set
Top 18%
1.4%
16
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.2%
17
Nature Computational Science
55 papers in training set
Top 1%
0.9%
18
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.9%
19
Journal of Cheminformatics
29 papers in training set
Top 0.7%
0.9%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 43%
0.6%
21
GigaScience
212 papers in training set
Top 5%
0.6%
22
BMC Bioinformatics
457 papers in training set
Top 6%
0.5%
23
Journal of Structural Biology
64 papers in training set
Top 1.0%
0.5%