Back

Decomposing metabolite set activity levels with PALS

McLuskey, K.; Wandy, J.; Vincent, I.; van der Hooft, J. J. J.; Rogers, S.; Burgess, K.; Daly, R.

2020-06-08 bioinformatics
10.1101/2020.06.07.138974 bioRxiv
Show abstract

MotivationRelated metabolites can be grouped into metabolite sets in many ways. Examples of these include the grouping of metabolites through their participation in a series of chemical reactions (forming metabolic pathways); or based on fragmentation spectral similarities and shared chemical substructures. Understanding how such metabolite sets change across samples can be incredibly useful in the interpretation and understanding of complex metabolomics data. However many of the available tools suitable for the enrichment analysis of metabolite sets are based on simple methods that badly handle the missing features inherent in untargeted metabolomics measurements and can be difficult to integrate into existing applications. ResultsWe present PALS (Pathway Activity Level Scoring), a Python library, command-line tool and Web application that performs the ranking of significantly-changing metabolite sets over different experimental conditions. As example applications, PALS is used to analyse metabolites grouped as pathways and by common MS-MS fragmentation structures. A comparison of PALS with two other commonly used methods (ORA and GSEA) is also given, and reveals that PALS is more robust to missing peaks and noisy data than the alternatives. We report results from using PALS to analyse pathways from a study of Human African Trypanosomiasis. Finally, we also report how PALS used tandem MS fragmentation structures to reveal enriched metabolite sets between clades in Rhamnaceae plant data, and on American Gut Project data. AvailabilityPALS is freely available from our project Web site at https://pals.glasgowcompbio.org/. It can be imported as a Python library, run as a stand-alone tool or used as a web application.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.