Orthrus: an AI-powered, cloud-ready, and open-source hybrid approach for metaproteomics
Chiang, Y.; Collins, M. J.
Show abstract
While metaproteomics provides invaluable insight into microbial communities and functions, significant bioinformatics challenges persist due to data complexity and the limitations of database searching. We introduce Orthrus, a hybrid approach combining transformer-based de novo sequencing (Casanovo) and database searching with rescoring (Sage+Mokapot). Benchmarking against PEAKS(R)11, MaxQuant, and MetaNovo, Orthrus demonstrates high peptide outputs, taxonomic diversity, and proteome coverage. Orthrus is Python-based and accessible to all via Google Colaboratory.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis and visualization of quantitative proteomics data using FragPipe-Analyst 96%
- TaxIt: An iterative and automated computational pipeline for untargeted strain-level identification using MS/MS spectra from pathogenic samples 96%
- The E. coli PeptideAtlas Build: Characterizing the observed Escherichia coli pan-proteome and its post-translational modifications 95%
Similar papers in this journal
- Critical Assessment of Metaproteome Investigation (CAMPI): a Multi-Lab Comparison of Established Workflows 96%
- AlphaPeptDeep: A modular deep learning framework to predict peptide properties for proteomics 96%
- Sequence-to-sequence translation from mass spectra to peptides with a transformer model 96%
Similar papers in this journal
- Target-Decoy MineR for determining the biological relevance of variables in noisy data sets 95%
- MS2AI: Automated repurposing of public peptide LC-MS data for machine learning applications 95%
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 94%
Similar papers in this journal
- An interactive mass spectrometry atlas of histone posttranslational modifications in T-cell acute leukemia 95%
- Implementing the re-use of public DIA proteomics datasets: from the PRIDE database to Expression Atlas 93%
- A peptide-centric quantitative proteomics dataset for the phenotypic assessment of Alzheimer's disease 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.