Quantifying biomarker ambiguity using metabolic network analysis
Hinkston, M. A.; Bradley, A. S.
Show abstract
Molecular biomarkers preserved in rocks provide evidence about ancient life but interpreting them requires inference through multiple stages of information loss arising from phylogenetic, biosynthetic, and diagenetic ambiguity. However, biomarker specificity is typically assessed qualitatively rather than quantitatively. Here we formalize biosynthetic ambiguity as entropy over metabolic networks. We introduce three metrics that quantify pathway-level information content: retrobiosynthetic complexity ({psi}), normalized branch depth ({lambda}), and fraction shared ({sigma}). Analysis of 9,140 MetaCyc metabolites defines a three-dimensional specificity space for biomarker evaluation. Only 13% of multi-pathway compounds exhibited low complexity, distal divergence, and high pathway consensus. Lipid biomarkers span this specificity space heterogeneously: hopanoids cluster near the high-specificity region while sterols occupy intermediate territory. Diagnostic quality and lipophilicity are approximately independent, so the constraint on molecular paleontology is the limited chemical diversity among preservable compound classes rather than their biosynthetic properties. This framework supports probabilistic biomarker interpretation by explicitly incorporating biosynthetic, phylogenetic, and diagenetic constraints.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Longitudinal wastewater sampling in buildings reveals temporal dynamics of metabolites 93%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 92%
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.