Back

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome

Chen, C.; Vakirlis, N.; Holmer, R.; de Ridder, D.; Kupczok, A.

2026-01-17 genomics
10.64898/2026.01.16.699869 bioRxiv
Show abstract

Orphan genes - genes lacking detectable homologs outside a species - are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted orphan genes may represent novel functional coding sequences. False positive orphan genes, also called spurious orphan genes, can arise from gene prediction errors. We reason that orphan genes lacking detectable expression are more likely to be spurious. To this end, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed orphan genes from spurious ones and to compare them with conserved genes found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified [~]218,000 orphan genes supported by expression evidence, while [~]330,000 predicted orphan genes lacked detectable expression, and were classified as spurious. We extracted 154 sequence, structural, and evolutionary features for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed orphan genes from spurious orphan genes and 0.93 in distinguishing expressed orphan genes from conserved genes. SHAP-based interpretation revealed clear biological signals. E.g., expressed orphans were present in more genomes than spurious ones and expressed orphan genes were shorter than conserved genes. This work improves orphan gene discovery and suggests that expressed orphan genes differ systematically from conserved genes and spurious orphan genes in sequence composition, structural constraints, and evolutionary signals.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.