Towards a Dataset for State of the Art Protein Toxin Classification
Challacombe, C. A.; Haas, N. S.
Show abstract
In-silico toxin classification assists in industry and academic endeavors and is critical for biosecurity. For instance, proteins and peptides hold promise as therapeutics for a myriad of conditions, and screening these biomolecules for toxicity is a necessary component of synthesis. Additionally, with the expanding scope of biological design tools, improved toxin classification is essential for mitigating dual-use risks. Here, a general toxin classifier that is capable of addressing these demands is developed. Applications for in-silico toxin classification are discussed, conventional and contemporary methods are reviewed, and criteria defining current needs for general toxin classification are introduced. As contemporary methods and their datasets only partially satisfy these criteria, a comprehensive approach to toxin classification is proposed that consists of training and validating a single sequence classifier, BioLMTox, on an improved dataset that unifies current datasets to align with the criteria. The resulting benchmark dataset eliminates ambiguously labeled sequences and allows for direct comparison against nine previous methods. Using this comprehensive dataset, a simple fine-tuning approach with ESM-2 was employed to train BioLMTox, resulting in accuracy and recall validation metrics of 0.964 and 0.984, respectively. This LLM-based model does not use traditional alignment methods and is capable of identifying toxins of various sequence lengths from multiple domains of life in sub-second time frames.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
- PhANNs, a fast and accurate tool and web server to classify phage structural proteins 93%
- Designing diverse and high-performance proteins with a large language model in the loop 93%
Similar papers in this journal
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 94%
- Protein language models can capture protein quaternary state 94%
- Application of a Machine Learning Approach Towards the Targeted Identification of Phage Depolymerases 93%
Similar papers in this journal
- Designing of thermostable proteins with a desired melting temperature 94%
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 93%
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.