Protein and genomic language models chart a vast landscape of antiphage defenses
Mordret, E.; Herve, A.; Vaysset, H.; Clabby, T.; Tesson, F.; Shomar, H.; Lavenir, R.; Cury, J.; Bernheim, A.
Show abstract
The bacterial pangenome encodes an immense array of antiphage systems, yet much of their diversity remains uncharted. In this study, we developed language models to predict novel antiphage proteins in two ways: first via fine-tuning ESM2, a protein language model capable of detecting distant homology to known defense proteins, second via a genomic language model with ALBERT architecture which predicts defensive function based on genomic context. We demonstrate that applying these approaches to Actinomycetota - a phylum largely unexplored for antiphage defenses, can accurately predict previously unknown functional defense mechanisms, leading to the discovery and experimental validation of six defense systems with novel antiphage proteins. Analysis of over 30,000 bacterial genomes predicted more than 45,000 uncharacterized protein families potentially involved in antiphage defense, underscoring the vast, untapped diversity of these systems.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SHARK enables homology assessment in unalignable anddisordered sequences 98%
- Versatile NTP recognition and domain fusions expand the functional repertoire of the ParB-CTPase fold beyond chromosome segregation 96%
- Revealing 29 sets of independently modulated genes in Staphylococcus aureus, their regulators and role in key physiological responses 96%
Similar papers in this journal
- Single-cell copy number lineage tracing enabling gene discovery 96%
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 96%
- CREaTor: zero-shot cis-regulatory pattern modeling with attention mechanisms 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.