Back

PDBCleanV2: A Python Library for Generating Consistent Structure Datasets

Pardo-Avila, F.; Weiner, L.; Cabral, P.; Corsepius, N. C.; Poitevin, F.; Levitt, M.

2025-02-19 bioinformatics
10.1101/2025.02.14.638326 bioRxiv
Show abstract

The number of structures in the Protein Data Bank has grown rapidly in recent years. Nonetheless, comparing structures of a specific system often proves challenging due to variant nomenclature, origins from different species, presence of various ligands, and missing atoms or chains. To address these issues, we have developed PDBCleanV2, a Python package that enables users to create consistent datasets of structures, simplifying comparison among structures. This library generates individual files for each molecule in a structure file, corrects labeling errors, and standardizes chain names and numbering. Our aim is to provide researchers with consistent datasets that streamline their analysis. The source code, installation instructions, and tutorials can be found in PDBCleanV2s GitHub repository and Zenodo accessible at https://github.com/fatipardo/PDBClean-0.0.2/ and https://doi.org/10.5281/zenodo.14014241, respectively.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.