Back

Proteogenomic discovery of novel small proteins in clinical Mycobacterium tuberculosis strains

Heiniger, B.; Schori, C.; Arefian, M.; Banaei-Esfahani, A.; Schuler, M.; Borrell Farnov, S.; Loiseau, C.; Brites, D.; Comas Espadas, I.; Aebersold, R.; Gagneux, S.; Collins, B. C.; Ahrens, C. H.

2026-01-27 microbiology
10.64898/2026.01.27.701740 bioRxiv
Show abstract

Even though our meta-analysis ranks Mycobacterium tuberculosis genomes among the bacterial pathogens that are most straightforward to assemble, most available assemblies rely on short-read sequencing and contain genomic blind spots that miss functionally important genes. Complete genomes are essential for functional genomics, particularly for identifying small ORF-encoded proteins (SEPs; [≤]100 amino acids), which can play critical biological roles yet are frequently missed by standard annotations. Here, we generated complete long-read assemblies for six clinical reference strains representing lineage 1 and the more pathogenic lineage 2, followed by comparative genomic and proteogenomic analyses. We additionally provide software to predict comprehensive sets of mycobacteria-specific PE and PPE family genes, including lineage-specific variants. Using parallel accumulation-serial fragmentation mass spectrometry, we detected approximately two-thirds of each strains annotated proteome from unfractionated cell extracts. Extending our proteogenomic framework across related strains, and rigorously controlling proteogenomic discovery rates using entrapment strategies, we revealed 12-24 previously unannotated proteins per strain, predominantly SEPs, 56-60 alternative translation start sites, and 9-17 expressed pseudogenes. Newly identified proteins included conserved and lineage-specific SEPs, an antitoxin, candidate antimicrobial peptides, and novel proteins under purifying selection. Overall, applying proteogenomics to phylogenomically selected clinical reference strains provides a valuable approach for discovering candidate diagnostics or therapeutics, as illustrated here for a WHO-listed critical bacterial pathogen.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.