Inferring Gene Presence in Incomplete Data via Phylogenetic Occupancy Modeling
Mattick, J. S. A.; DeMontigny, W. C.; Delwiche, C. F.
Show abstract
Increasing access to genomic data has revolutionized our understanding of biology. Organisms that were previously unculturable or otherwise difficult to study have been investigated using metagenomic sequencing and bioinformatic assemblies, illuminating biological diversity that was previously invisible. However, as the availability of genomic data has grown, so has the challenge posed by incomplete genomes. Many genomes obtained from metagenomic assemblies or mixed cultures are of poor quality and establishing fully complete genomes requires substantial effort. Incomplete genomes pose difficulties for several common analyses, particularly gene-inventory and core-genome analyses. When genomes are incomplete, distinguishing true gene absence from non-detection becomes difficult. For relatively complete genomes, gene absences are inferred to be true absences, but for highly incomplete genomes, researchers often exclude such data entirely. Probabilistic models have attempted to address this issue in core genome analyses. In particular, mO-TUpan utilizes an iterative algorithm to infer genome completeness and categorize genes as core or accessory. In this work, we substantially improve upon this approach by integrating a well-established class of ecological models, called occupancy models, with evolutionary modeling. Our "phylogenetic occupancy model" defines a probability distribution over gene presence that accounts for shared information across related genomes. This framework simultaneously estimates genome completeness and the probability that a gene is present but unobserved. This model substantially outperforms competing methods for core genome inference and enables inference of single-gene presence/absence and ancestral-state reconstruction. Alongside this paper, we provide our model as a Python package.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.