Metabolites
○ MDPI AG
All preprints, ranked by how well they match Metabolites's content profile, based on 53 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Choudhary, K. S.; Fahy, E.; Coakley, K.; Sud, M.; Maurya, M. R.; Subramaniam, S.
Show abstract
With the advent of high throughput mass spectrometric methods, metabolomics has emerged as an essential area of research in biomedicine with the potential to provide deep biological insights into normal and diseased functions in physiology. However, to achieve the potential offered by metabolomics measures, there is a need for biologist-friendly integrative analysis tools that can transform data into mechanisms that relate to phenotypes. Here, we describe MetENP, an R package, and a user-friendly web application deployed at the Metabolomics Workbench site extending the metabolomics enrichment analysis to include species-specific pathway analysis, pathway enrichment scores, gene-enzyme information, and enzymatic activities of the significantly altered metabolites. MetENP provides a highly customizable workflow through various user-specified options and includes support for all metabolite species with available KEGG pathways. MetENPweb is a web application for calculating metabolite and pathway enrichment analysis. Availability and ImplementationThe MetENP package is freely available from Metabolomics Workbench GitHub: (https://github.com/metabolomicsworkbench/MetENP), the web application, is freely available at (https://www.metabolomicsworkbench.org/data/analyze.php)
Huang, L.; Preciat, G.; Alarcon-Gil, J.; Moreno, E. L.; Wegrzyn, A. B.; Thiele, I.; Schymanski, E. L.; Harms, A. C.; Fleming, R. M.; Hankemeier, T.
Show abstract
Quantitative inference of intracellular reaction rates is essential for characterising metabolic phenotypes. The classical experimental method for measuring metabolic fluxes makes use of stable-isotope tracing of metabolites through the metabolic network, followed by mass spectrometry analysis. The most common 13C-based metabolic flux analysis requires multidisciplinary knowledge in analytical chemistry, cell biology, and mathematical modelling, as well as the use of multiple independent tools for handling mass spectrometry data. Besides, flux analysis is usually carried out within a small network to validate a specific biological hypothesis. To overcome interdisciplinary barriers and extend flux interpretation towards a genome-scale level, we developed fluxTrAM, a semi-automated pipeline for processing tracer- based metabolomics data and integrating it with atomically resolved genome-scale metabolic networks to enable flux predictions at genome-scale. fluxTrAM integrates different software packages inside and outside of the COBRA Toolbox v3.4 for the generation of metabolite structure and reaction databases for a genome-scale model, labelled mass spectrometry data processing into standardised mass isotopologue distribution data (MID), and metabolic flux analysis. To demonstrate the utility of this pipeline, we generated 13C-labeled metabolomics data on an in vitro human induced pluripotent stem cell (iPSC)-derived dopaminergic neuronal culture and processed 13C-labeled MID datasets. In parallel, we generated a cheminformatic database of standardised and context-specific metabolite structures, and atom-mapped reactions for a genome-scale dopaminergic neuronal metabolic model. MID data could be exported into established flux inference software for conventional flux inference on a core model scale. It could also be integrated into the atomically resolved metabolic model for flux inference at genome-scale using moiety fluxomics method. The core model flux solution and moiety flux solution were then compared to two additional flux solutions predicted via flux balance analysis and entropic flux balance analysis. The extensive computational flux analysis and comparison helped to better evaluate the obtained flux feasibility of the neuron-specific genome-scale model and suggested new tracer-based metabolomics experiments with novel labeling configurations, such as labelling a moiety within the thymidine metabolite. Overall, fluxTrAM enables the automation of labelled liquid chromatography (LC)-mass spectrometry (MS) data processing into MID datasets and atom mapping for any given genome-scale metabolic model. It contributes to the standardisation and high throughput of metabolic flux analysis at genome- scale.
Godlewski, A.; Solowiej, K.; Mojsak, P.; Godzien, J.; Zelkowska, J.; Kretowski, A.; Lyson, T.; Burdukiewicz, M.; Kaminski, K.; Ciborowski, M.
Show abstract
Class imbalance remains a challenge in metabolomics research, where biological and technical variability can affect statistical inference and machine learning (ML) performance. Class-balancing algorithms address this issue by either increasing minority-class observations or reducing the number of majority-class samples. This study evaluated the impact of oversampling and undersampling algorithms on targeted and untargeted metabolomics datasets derived from LC-MS and GC-MS analyses of plasma samples from patients with glioblastoma, meningioma, and controls. Synthetic Minority Oversampling Technique (SMOTE) and Random Undersampling (RUS) were applied to balance the datasets, and their effects on data distribution, inter-feature correlations, and machine learning model performance were compared. RUS preserved the original feature distributions but reduced representativeness by removing the majority-class samples. In contrast, SMOTE introduced synthetic samples that altered covariance structures, increasing the risk of overfitting, particularly in small datasets (n=10). These effects diminished with larger groups (n=30), partially restoring correlations between metabolites. Model performance varied across the class-balancing algorithms. Random Forest classifiers benefited from both balancing methods, with undersampling often yielding higher F1 scores, whereas Support Vector Machine models showed reduced classification performance. These findings highlight the importance of selecting class-balancing strategies based on dataset size, analytical platform, and ML algorithm in metabolomics studies.
Mekonnen, Y. A.; Dhake, N.; Jaiswal, S.; Rubio, V.; Narvaez-Bandera, I.; Ackerman, H.; Flore, E.; Stewart, P.
Show abstract
Advances in metabolomics have significantly improved our understanding of cellular processes by enabling the identification of hundreds of metabolites in a single experiment. These developments provide valuable insights into complex metabolic networks. While efforts have been made to develop pathway enrichment analysis (PEA), existing implementation often require multiple steps, rely on web-based interfaces, or depend on R packages configuration that may affect reproducibility and ease of use. To overcome these limitations, we introduce EnrichMet, an R package for fast, flexible, and reproducible pathway enrichment analysis. EnrichMet modules support over-representation analysis of pathways, metabolite set enrichment analysis (MetSEA), and network-based pathway analysis. The package streamlines the workflow by combining curated pathway information from the Kyoto Encyclopedia of Genes and Genomics (KEGG) and employs Fishers Exact Test to identify significantly enriched pathways. Benchmark analyses show that enrichment on sample data completes in approximately 3 seconds. EnrichMet offers both a command-line and a user-friendly Shiny interface, enabling accessibility for users with or without programming experience. Through case studies on experimental metabolomics datasets, we demonstrated that EnrichMet delivers accurate and comprehensive pathway enrichment results while minimizing computational time and simplifying user interaction. Furthermore, its flexible framework supports extensions to other data types and knowledge bases beyond KEGG, as illustrated through a lipidomics case study. By unifying performance, reproducibility, usability, and visualization within a single package, EnrichMet facilitates deeper insights and promotes efficient, transparent, and reproducible research practices. Availability and implementation(https://github.com/biodatalab/enrichmet.git)
Smith, A.; Pinto, R.; Zagkos, L.; Tzoulaki, I.; Elliott, P.; Dehghan, A.
Show abstract
BackgroundMetabolomics data are often generated through different analytical platforms and different methods of identification and quantification which makes their synthesis and large-scale replication challenging. To address this, we applied generative deep learning to impute metabolites assayed by Metabolon, a commonly used commercial platform, using metabolomic features acquired by an untargeted liquid chromatography-mass spectrometry (LC-MS) platform. MethodsWe utilised a subset of 979 samples from the Airwave Health Monitoring Study which were assayed by both Metabolon and National Phenome Centre at Imperial College (NPC) LC-MS assays to develop an ensemble of importance-weighted autoencoders (IWAEs) which can perform cross-platform metabolomics imputation between the two assays. Using the ensemble, we generated a Metabolon equivalent dataset in 2,971 additional Airwave samples that lacked prior Metabolon measurements. We conducted observational associations with two clinical outcomes, body mass index (BMI) and C-reactive protein (CRP). We validated the ensemble and imputed data by investigating the concordance of the observational associations. This was done using both the imputed Metabolon dataset and the measured metabolite levels by Metabolon, and NPC in the Airwave study and Nightingale platform in the UK Biobank. ResultsOur imputation ensemble generated samples highly correlated with their real values across all Metabolon metabolites within a held-out test set with a mean sample correlation of 0.61 (IQR 0.55-0.67). The well-imputed subset included 199 (22%) of the metabolites present in the real Metabolon dataset where the imputed values accounted for at least 55% of the original variance (R2 [≥] 0.55) and a minimal uncertainty (R2 variance [≤] 0.025). The subset included 43 metabolites not previously identified within our LC-MS platform. When comparing the associations of the real and imputed Metabolon metabolites with BMI and CRP, the standardised beta-coefficients were highly correlated ({rho} = 0.93 for BMI and 0.89 for CRP) with minimal mean difference (0.005 (0.04) for BMI, 0.005 (0.04) for CRP). Similar concordance occurred between the imputed Metabolon metabolites and equivalent UK Biobank (mean difference -0.007 (0.05) for BMI, 0.01 (0.04) for CRP) and our LC-MS platform (mean difference -0.013 (0.04) for BMI, -0.019 (0.04) for CRP). ConclusionThis methodological innovation offers a scalable and accurate method for cross-platform imputation which could allow for to aggregate individual-level metabolomics data from different epidemiological studies, replication findings or conduct meta-analyses.
Niazi, U.; Roberts, C. A.; McDonnell, D.; Goss, V. M.; Afolabi, P. R.; Swann, J. R.; Byrne, C. D.; Griffiths, G. O.; Hamady, Z. Z.
Show abstract
Background: Early detection of pancreatic ductal adenocarcinoma (PDAC) is critical. While faecal elastase-1 (FE-1) is a standard clinical marker for pancreatic function, its diagnostic accuracy for malignancy is limited. We sought to identify plasma metabolites that enhance FE-1 performance in symptomatic "at-risk" patients. Methods: Using the DEPEND cohort (CRUK C45617/A29908), plasma metabolomics was performed on patients with resectable PDAC (n=23) and healthy volunteers (n=24). Predictive modelling included feature selection and cross-validation, with further validation in an independent external cohort. Results: Citrulline was identified as significantly depleted in PDAC patients across discovery and validation cohorts. In isolation, Citrulline achieved an AUC of 0.86 (internal) and 0.88 (external validation). Standalone FE-1 demonstrated an AUC of 0.67. However, combining Citrulline and FE-1 significantly improved diagnostic performance, achieving a combined AUC of 0.96. Stratification revealed distinct metabolomic signatures associated with poorly differentiated tumours, suggesting a link to histological grade. Conclusions: Integrating Citrulline with FE-1 testing substantially improves PDAC detection in symptomatic patients. This non-invasive panel offers high diagnostic potential, though prospective validation is required to establish clinical cut-offs for routine practice.
Bambarandhage, A.; Zainurin, A. A.; Laziri, N.; Gate, T.; Tench, H.; Beckmann, M.; Phillips, H.; Morphew, R.; Pennick, M. O.; Mur, L. A.
Show abstract
IntroductionBreast Cancer (BC) remains a significant clinical challenge, and despite well-established screening strategies, new biomarkers could improve BC detection, treatment and management. Urine represents a headily accessible liquid biopsy for diagnosis and extracellular vesicle (EV) transfer of oncogenic proteins, RNAs, and metabolites that promote tumor growth, invasion, metastasis, and immune evasion. AimsTo compare the whole urine and urinary EV metabolomes and identify BC specific metabolite changes. MethodologyUrine samples were collected from four participant groups: breast cancer (BC) patients (n = 42), individuals with breast benign disease (BBD; n = 3), symptom controls (SC; n = 4), and healthy controls (HC; n = 6). EVs were isolated using differential centrifugation, ultrafiltration, and size-exclusion chromatography (SEC), and their morphology was confirmed by transmission electron microscopy (TEM). Metabolites from whole urine and from EVs derived from the same samples were extracted using methanol-water (70:30, v/v) and analyzed by direct-infusion mass spectrometry (DI-MS) in both positive and negative ESI modes. Metabolic features were processed with BinneR and annotated using the HMDB and KEGG databases. Integrated multi-omics analysis of whole-urine and EV-associated metabolomes was performed using the DIABLO framework within the MixOmics package in R platform. ResultsDI-MS profiling detected a broad spectrum of metabolites in both whole-urine and EV-derived fractions. Multivariate analyses revealed a clear separation of breast cancer (BC) patients from healthy controls and non-cancer groups in both matrices. Whole EV metabolites with area under the curves (AUC) of > 0.7 included glyceryl phosphoryl derivatives, N-eicosapentaenoyl species, sphinganine-1-phosphate and tetracosahexaenoic acid. EV-enriched metabolites included carnitine, histidine and adenosine monophosphate. DIABLO-based integrative analysis suggested that urinary and EV metabolomes were broadly similar with the discrete putative metabolite biomarkers representing minor, but specific changes with BC. ConclusionsThe whole urine and EV metabolomes suggested a small number of metabolite changes that were specific to BC. This could indicate that the urinary EVs describe distinctive aspects of the breast carcinogenic process.
Muller, C.; Audemard, J.; Prigent, S.; Frioux, C.
Show abstract
MotivationMetabolic networks represent genome-derived information about the biochemical reactions that cells are capable of performing. Mapping omic data onto these networks is important to refine model simulations. However, metabolomic data mapping remains very challenging due to difficulties in identifier reconciliation between annotation profiles and metabolic networks. ResultsMetaNetMap is a Python package designed to automatise the process of mapping metabolomic data onto metabolic networks. It includes several layers of identifier matching, the use of customisable databases, and molecular ontology integration to suggest the most matches between experimentally-identified metabolites and molecules defined in the network. We demonstrate its usability and the quality of automated mapping using two datasets. Availability and ImplementationMetaNetMap is an open source python package and publicly available under the GPLv3 licence. Source code is freely available on GitHub: https://github.com/coraliemuller/metanetmap. Data and code used in the application cases of this paper can be found at: https://doi.org/10.57745/ESFLR8.
Hogg, M.; Wolfschmitt, E.-M.; Wachter, U.; Zink, F.; Radermacher, P.; Vogt, J. A.
Show abstract
The pentose phosphate pathway (PPP) plays a key role in the cellular regulation of immune cell function; however, little is known about the interplay of metabolic adjustments in granulocytes, especially regarding the non-oxidative PPP. For the determination of metabolic mechanisms within glucose metabolism, we propose a novel Bayesian 13C-Metabolic flux analysis based on ex-vivo parallel tracer experiments with [1,2-13C]glucose, [U-13C]glucose, and [4,5,6-13C]glucose and gas chromatography-mass spectrometry labeling measurements of metabolic fragments including sugar phosphates. With this approach we obtained precise flux distributions and their joint confidence regions, which showed that phagocytic stimulation reversed the direction of non-oxidative PPP net fluxes from ribose-5-phosphate biosynthesis towards glycolytic pathways. This process was closely associated with the up-regulation of the oxidative PPP to promote the oxidative burst. The estimated fluxes showed strong pairwise inter-relations forming a single line in several cases. This behavior could be explained with a three-dimensional permissible space derived from stoichiometric-flux-constraint analysis and enabled a principal component analysis detecting only three distinct axes of coordinated flux changes that were sufficient to explain all flux observations.
Cai, L.; Hieu, V.; Gu, W.; Chen, H.; Franklin, J.; Abou Haidar, L.; Wu, Z.; Pan, C.; Cai, F.; Nguyen, P.; Ko, B.; Yang, C.; Zacharias, L. G.; Sudderth, J.; Montgomery, S.; Uhles, C.; Fisher, H.; Hudnall, J.; Hornbuckle, C.; Quinn, C.; Michel, D.; Umana, L.; Scheuerle, A.; McNutt, M.; Gotway, G.; Afroze, B.; Ni, M.; DeBerardinis, R. J.
Show abstract
Metabolomic profiling is instrumental in understanding the systemic and cellular impact of inborn errors of metabolism (IEMs), monogenic disorders caused by pathogenic genomic variants in genes involved in metabolism. This study encompasses untargeted metabolomics analysis of plasma from 474 individuals and fibroblasts from 67 subjects, incorporating healthy controls, patients with 65 different monogenic diseases, and numerous undiagnosed cases. We introduce a web application designed for the in-depth exploration of this extensive metabolomics database. The application offers a user-friendly interface for data review, download, and detailed analysis of metabolic deviations linked to IEMs at the level of individual patients or groups of patients with the same diagnosis. It also provides interactive tools for investigating metabolic relationships and offers comparative analyses of plasma and fibroblast profiles. This tool emphasizes the metabolic interplay within and across biological matrices, enriching our understanding of metabolic regulation in health and disease. As a resource, the application provides broad utility in research, offering novel insights into metabolic pathways and their alterations in various disorders.
Huckvale, E. D.; Thompson, P. T.; Flight, R. M.; Moseley, H. N. B.
Show abstract
Background/ObjectivesMetabolism-level interpretation of metabolomics datasets requires aggregation analyses across metabolites. One highlyused aggregation analysis is pathway enrichment analysis (PEA), which involves detecting pathways enriched with metabolites that are differential between experimental groups. Annotating metabolites with pathway associations is a prerequisite for PEA. While several knowledgebases define pathways and include metabolite-pathway annotations, these definitions are often partially or even grossly incomplete due to limitations in current metabolic knowledge and its curation, which greatly limits the effectiveness of PEA. MethodsIn this work, we used a novel multitask classification, graph convolutional-like neural network to generate high-quality metabolite-pathway annotations for pathways defined across KEGG, MetaCyc, and Reactome. We then included these predicted metabolite-pathway annotations when performing PEA on 990 datasets deposited in Metabolomics Workbench. ResultsWe demonstrate an 8-fold increase in the median number of enriched pathways detected across these datasets compared to using only knowledgebase-derived annotations. ConclusionsThe significant increase in enriched pathways substantially improves the biological and biomedical interpretability of metabolomics datasets.
Flammer, E.; Higdon, L.; Sanda, S.; Garrett, T.; Ismail, H. M.
Show abstract
Aims/hypothesisImmunotherapies such as Teplizumab can preserve residual beta cell function in individuals with newly diagnosed type 1 diabetes (T1D), but treatment response is variable. Currently, no biomarker exists to identify individuals most likely to benefit from immunotherapy. We believe that baseline serum metabolomic profiles can distinguish individuals who respond to treatment from nonresponders and predict therapeutic response. MethodsBaseline serum samples from 41 individuals newly diagnosed with T1D enrolled in the AbATE trial (NCT00129259) were analyzed to identify metabolic predictors of response to Teplizumab therapy in the AbATE trial. Responders to Teplizumab, as per study protocol, were defined as individuals who exhibited less than a 40% decline in baseline C-peptide levels at 2 years after start of treatment. We analyzed baseline serum samples using a semi-targeted metabolomics approach via liquid chromatography-high-resolution tandem mass spectrometry. Metabolites that were significantly different between responders and nonresponders were identified (P < 0.05), and the significant metabolites were used to train a supervised Random Forest model to predict treatment response. Model performance was evaluated using a 70/30 training/testing split, 5-fold cross-validation, bootstrap resampling (1,000 iterations), and permutation testing (1,000 permutations). ResultsWe identified 15 significantly different metabolites at baseline between responders and nonresponders (P < 0.05). These metabolites included amino acids and their derivatives, tricarboxylic acid (TCA) cycle intermediates, and microbially derived metabolites. At baseline, responders exhibited higher levels of TCA cycle metabolites, amino acid derivatives, and microbial metabolites, whereas nonresponders showed elevated levels of glutamate and acylcarnitines. The Random Forest classifier achieved an accuracy of 0.769 and an area under the receiver operating characteristic curve (AUC) of 0.881 in the test dataset. Cross-validation yielded a mean AUC of 0.856 (SD 0.156; 95% CI 0.719-0.992). Bootstrap analysis produced a test AUC 95% CI of 0.619-1.000, and permutation testing confirmed significance (p = 0.012). Conclusions/interpretationBaseline serum metabolomic signatures can predict responders to Teplizumab with high accuracy. This could potentially be applicable when considering other immunotherapies in preventative efforts in T1D. Trial registrationClinicalTrials.gov NCT00129259. Research in ContextO_ST_ABSWhat is already known about this subject?C_ST_ABSO_LITeplizumab can delay beta cell decline in individuals with newly diagnosed T1D, but treatment response varies. C_LIO_LINo validated biomarkers currently exist to predict which individuals will respond to immunotherapy. C_LIO_LIMetabolomic profiling has shown potential for identifying metabolic signatures associated with disease progression and immune activity in T1D. C_LI What is the key question?O_LICan baseline serum metabolomic profiles predict which individuals with newly diagnosed T1D will respond to Teplizumab therapy? C_LI What are the new findings?O_LIFifteen baseline metabolites differed significantly between responders and nonresponders, including amino acid derivatives, tricarboxylic acid cycle intermediates, and microbially derived metabolites. C_LIO_LIResponders exhibited metabolic signatures consistent with preserved beta cell function and enhanced mitochondrial and immune-regulatory activity. C_LIO_LIA Random Forest model developed using these metabolites accurately predicted treatment response (AUC 0.881), demonstrating strong predictive potential. C_LI How might this impact on clinical practice in the foreseeable future?O_LIBaseline metabolomic profiling could support personalized treatment strategies by identifying individuals most likely to benefit from treatment with Teplizumab or other immunotherapies. C_LI
Smith, M. l.; Goudswaard, L. J.; Hughes, D. A.; Blazeby, J. M.; Rogers, C. A.; Mazza, G.; Gidman, E. A.; Fitzgibbon, S.; Groom, A.; Ring, S. M.; Timpson, N. J.; Corbin, L. J.
Show abstract
Metabolomics data has been generated via proton nuclear magnetic resonance (NMR) spectroscopy in samples collected within By-Band-Sleeve. Two sample collection efforts were made - firstly, from a randomised controlled trial (RCT) comparing the effectiveness of three types of bariatric surgery: the Roux-en-Y gastric bypass ("bypass"), laparoscopic adjustable gastric band ("band") and the sleeve gastrectomy ("sleeve"), and secondly from a non-randomised (observational) study of bariatric surgery. In both instances, samples were collected from patients before and after surgery. Data underwent quality control (QC) using a standard pipeline via the R package metaboprep. This package extracts data from preformed worksheets, provides summary statistics and enables the user to select samples and metabolites for their analysis based on a set of quality metrics. Post-filtering, the dataset consists of data from 1410 samples (999 pre-surgery, 411 post-surgery) from 1045 unique individuals (1000 from the RCT and 45 from the non-randomised study), each with 250 measured metabolic traits. Comparison of NMR measures to clinical chemistry data showed good agreement for the metabolites in common across both datasets. Concordance with previous NMR data generated for a subset of the same samples was largely good. Overall, this data note describes the data, explains the pre-processing and quality control procedures applied to the data, and provides some data validation analyses.
Anctil, N.; Hauguel, P.; Noel, L.-P.
Show abstract
BackgroundBreast cancer (BC) remains the most diagnosed malignancy and leading cancer-related cause of mortality in women worldwide. Although blood-based untargeted metabolomics has emerged as a promising modality for detecting early-stage BC, the clinical translation of this approach has been bottlenecked by two unresolved issues: (i) the field has almost exclusively relied on serum or plasma, which require venipuncture and cold-chain logistics, and (ii) machine-learning models reported on such data are frequently validated with protocols that are blind to analytical batch structure, producing optimistically biased performance estimates. MethodsWe present a breast cancer detection study based on dried blood spots (DBS), an analytical matrix that enables self-collection and ambient-temperature shipping. A cohort of 2,734 participants (114 biopsy-confirmed BC cases; 2,620 non-cancer controls) was profiled by untargeted LC-MS/MS on a Thermo Scientific Orbitrap IQ-X coupled to a Vanquish UHPLC. A 39-metabolite panel meeting MSI Level 1 identification criteria [1] was pre-specified a priori from the published breast-cancer metabolomics literature, frozen prior to LC-MS acquisition, and applied to the present cohort without any feature selection on the data. Six standard supervised-learning architectures (LASSO, Elastic Net, Linear SVM, PLS-DA, OPLS-DA, XGBoost) were evaluated on this pre-specified panel; OPLS-DA, whose pyopls implementation does not integrate cleanly into the repeated multi-seed batch-aware protocol, is reported only in the sex-matched subgroup analysis where a single-seed 5-fold stratified protocol permits a directly comparable fit. Per-batch control-median normalization is applied upstream, following the protocol of the companion same-lab study [2], which removes batch-specific intensity shifts at the data-preparation stage; kNN imputation, log transform, and robust scaling are then fit within each training fold. The evaluation battery comprises batch-aware StratifiedGroupKFold CV reported at single-seed (seed=42) with inter-seed SD quantified across 10 independent seeds, batch-aware nested CV, a 100-seed held-out 20%-batch validation with disjoint-batch isotonic probability calibration (30% calibration partition), PPV/NPV reporting at multiple operating points and three deployment prevalences, subgroup analyses by TNM stage and tumor grade, pathway-ablation sensitivity analysis, and a 1,000-iteration permutation test. ResultsUnder batch-aware evaluation (StratifiedGroupKFold, single-seed=42), AUC ranged from 0.914 to 0.949 across classifiers, with LASSO achieving 0.928 and XGBoost 0.949; inter-seed SD across 10 seeds was 0.002-0.006. At 95% specificity, LASSO reached 75.4% sensitivity and XGBoost 81.6%. Held-out batch validation (100 seeds) yielded mean AUC 0.912 for Elastic Net and 0.935 for XGBoost, confirming robust generalization. All 39 panel features showed high coefficient stability, and permutation testing on representative classifiers (LASSO, Linear SVM, PLS-DA) yielded p [≤] 0.001. Subgroup analyses showed weaker detection of stage IIA tumors (AUC 0.87, n=40) compared with stage IIB/IIIA (AUC 0.95), consistent with stronger metabolic signatures in more advanced disease. Bootstrap coefficient consistency of the Elastic Net classifier confirmed that all 39 panel features received a non-zero multivariate weight in >=80% of 100 stratified bootstraps. Permutation testing on the three representative classifiers subjected to this analysis (LASSO, Linear SVM, PLS-DA) confirmed significance at p [≤] 0.001 in all three cases. ConclusionsOn this cohort of diagnosed, pre-treatment breast-cancer cases, DBS LC-MS metabolomic profiling delivers classification performance (AUC 0.928 for LASSO and 0.949 for XGBoost under batch-aware GroupKFold CV at single-seed=42; held-out AUC 0.912-0.935) that is robust across classifier families and biological pathways. The DBS matrix is non-radiating, self-collectable by finger-prick, and mailable at ambient temperature. The approach complements the established venous-blood workflow while addressing a clear infrastructural gap identified over nearly a decade of preliminary work [3, 4]. Performance is weaker on stage IIA than on more advanced disease, and prospective validation in an independent asymptomatic screening cohort is required before clinical positioning as a decentralized triage modality.
Chocholouskova, M.; Ctvrtlik, F.; Tudos, Z.; Hartmann, I.; Schovanek, J.; Vostalova, J.; Proskova, J.; Pacak, K.; Holcapek, M.
Show abstract
Adrenocortical carcinoma (ACC) is a rare, aggressive malignancy posing significant diagnostic challenges, particularly in distinguishing it from other adrenal tumors, such as adenoma and pheochromocytoma, due to overlapping imaging and biochemical features. Improved non-invasive tools are critically needed for earlier, more accurate classification of this rare cancer. This pilot study analyzed serum lipidomic profiles in ACC, pheochromocytoma, and adenoma patients versus healthy volunteers. The most significant alterations occurred in sphingomyelins (SM) and diacylglycerols (DG). All tumor samples showed reduced very-long odd-chain SM (e.g., SM 39:1, SM 41:1, SM 41:2) and elevated DG (e.g., DG 34:1, DG 34:2, DG 36:2). These abnormalities were most pronounced in malignant tumors: ACC and metastases (AUC = 0.933), followed by pheochromocytoma (AUC = 0.800) and adenoma (AUC = 0.711). ACC patients also exhibited specific lipid signatures with decreased alkyl/alkenyl phospholipids (e.g., PE O-38:5) and lysophosphatidylcholines (e.g., LPC 20:5, LPC 18:2) versus healthy volunteers, not observed in pheochromocytoma or adenomas. Ceramide species (e.g., Cer 42:2;O2, Cer 34:1;O2) were increased in ACC compared to the other tumor types. Incorporating lipid-to-lipid ratios (Cer/SM, Cer/DG) further improved statistical model accuracy. Compared to clinical biochemistry/oxidative stress (OS) parameters, lipidomic profiling showed superior discriminatory power in adrenal tumor diagnosis. The presented study shows the serum lipidomic profiling as a promising non-invasive method for distinguishing adrenal tumor subtypes (ACC, pheochromocytoma, and adenoma) from healthy individuals, with strong diagnostic potential for ACC.
Zararsiz, G. E.; Lintelmann, J.; Cecil, A.; Kirwan, J.; Poschet, G.; Gegner, H. M.; Schuchardt, S.; Guan, X. L.; Saigusa, D.; Wishart, D.; Zheng, J.; Mandal, R.; Adams, K.; Thompson, J. W.; Snyder, M. P.; Contrepois, K.; Chen, S.; Ashrafi, N.; Akyol, S.; Yilmaz, A.; Graham, S. F.; O'Connell, T. M.; Kalecky, K.; Bottiglieri, T.; Limonciel, A.; Pham, H. T.; Koal, T.; Adamski, J.; Kastenmüller, G.
Show abstract
Metabolomics and lipidomics are pivotal in understanding phenotypic variations beyond genomics. However, quantification and comparability of mass spectrometry (MS)-derived data are challenging. Standardised assays can enhance data comparability, enabling applications in multi-center epidemiological and clinical studies. Here we evaluated the performance and reproducibility of the MxP(R) Quant 500 kit across 14 laboratories. The kit allows quantification of 634 different metabolites from 26 compound classes using triple quadrupole MS. Each laboratory analysed twelve samples, including human plasma and serum, lipaemic plasma, NIST SRM 1950, and mouse and rat plasma, in triplicates. 505 out of the 634 metabolites were measurable above the limit of detection in all laboratories, while eight metabolites were undetectable in our study. Out of the 505 metabolites, 412 were observed in both human and rodent samples. Overall, the kit exhibited high reproducibility with a median coefficient of variation (CV) of 14.3 %. CVs in NIST SRM 1950 reference plasma were below 25 % and 10 % for 494 and 138 metabolites, respectively. To facilitate further inspection of reproducibility for any compound, we provide detailed results from the in-depth evaluation of reproducibility across concentration ranges using Deming regression. Interlaboratory reproducibility was similar across sample types, with some species-, matrix-, and phenotype-specific differences due to variations in concentration ranges. Comparisons with previous studies on the performance of MS-based kits (including the AbsoluteIDQ p180 and the Lipidyzer) revealed good concordance of reproducibility results and measured absolute concentrations in NIST SRM 1950 for most metabolites, making the MxP(R) Quant 500 kit a relevant tool to apply metabolomics and lipidomics in multi-center studies.
Wood, R. M.; Corbin, L. J.; Blazeby, J. M.; Rogers, C.; Timpson, N. J.; Lawson, D. J.
Show abstract
High throughput metabolomic assays offer a huge opportunity to quantify the cellular processes underlying disease and intervention pathways. However, the multi-dimensional inter-relatedness between these processes coupled with the complex noisy measurement environment create a need for generation of new methods that move beyond simple pairwise associations. Here we develop a computationally simple, multivariate, relational comparison method called CLARITY to compare metabolomic data before and after an intervention. This generates a relational anomaly score that combines with traditional methods to increase classification performance of the underlying cause of changes to the levels of and covariances between metabolites. We demonstrate utility in the By-Band-Sleeve (BBS) clinical trial of bariatric surgery using NMR metabolomics data. On supplementing linear regression analysis with CLARITY, previously identified changes form two clusters that imply involvement in different underlying biological pathways. An additional cluster of metabolites are identified as undergoing a relational change which would not have been detected using traditional methods. Gathering insights about metabolites and the biomarkers they capture in the causal pathway between intervention and effect, from observations at scale, will inform the future design of modelling and laboratory experiments to capture the underlying biological process.
Krutkin, D. D.; Thomas, S.; Zuffa, S.; Rajkumar, P.; Knight, R.; Dorrestein, P. C.; Kelley, S. T.
Show abstract
Untargeted metabolomics often produce large datasets with missing values, arising from biological or technical factors, which can undermine statistical analyses and lead to biased biological interpretations. Imputation methods, such as k-Nearest Neighbors (kNN) and Random Forest (RF) regression are commonly used but their effects vary depending on the type of missing data e.g. Missing Completely At Random (MCAR) and Missing Not At Random (MNAR). Here, we determined the impacts of degree and type of missing data on the accuracy of kNN and RF imputation using two datasets: a targeted metabolomic dataset with spiked-in standards and an untargeted metabolomic dataset. We also assessed the effect of compositional data approaches (CoDA), such as the centered log-ratio (CLR) transform, on data interpretation, since these methods are increasingly being used in metabolomics. Overall, we found that kNN and RF performed more accurately when the proportion of missing data across samples for a metabolic feature was low. However, these imputations could not handle MNAR data and generated wildly inflated values or imputed values where none should exist. Furthermore, we show that the proportion of missing values had a strong impact on the accuracy of imputation which affected the interpretation of the results. Our results suggest extreme caution should be used with imputation even with modestly levels of missing data or when the type of missingness is unknown.
Rocha, B. L.; Jonaitis, E. M.; Hamwi, A.; Engelman, C. D.
Show abstract
Background/ObjectivesLongitudinal metabolomics analysis offers valuable insight into how metabolic pathways change according to age and health status. However, metabolite levels can fluctuate due to biological factors (ex. age, diet, health-status) and technical factors (ex. sample handling, storage times, instrument performance), with some metabolites exhibiting greater sensitivity to these sources of variability than others. This study aimed to characterize the longitudinal and technical stability of untargeted plasma and cerebrospinal fluid (CSF) metabolites, and to identify a subset that remains reliable over the extended time scales required for epidemiological research. MethodsUntargeted ultra-high-performance liquid chromatography-mass spectrometry (LC-MS) metabolomic profiles were available from multiple visits in the Wisconsin Registry for Alzheimers Prevention (WRAP) and Wisconsin Alzheimers Disease Research Center (ADRC) studies. For this analysis, we constructed a subset of generally healthy participants with samples drawn at four time points ([~]2.5 years apart): two visits analyzed in 2017 and two visits analyzed in 2023, corresponding to two distinct analytical waves. We computed Rotherys intraclass correlation coefficients (ICCs) to quantify intrawave and inter-wave stability, evaluated pooled quality-control (QC) variation, classified metabolite stability by established thresh-olds, and developed a composite score integrating longitudinal stability and susceptibility to technical variance. ResultsAcross all metabolites, median stability was classified as fair (Rotherys{rho} >0.40 to [≤]0.75) for both plasma and CSF. Although analytical batches were bridged using pooled QC samples, inter-wave stability was significantly lower than intra-wave stability, reflecting increased technical variability across waves. Using the composite score, we identified subsets of metabolites with excellent stability and low susceptibility to batch effects in plasma and CSF. Stability patterns varied across biochemical super pathways. ConclusionsThis work highlights metabolites suitable for long-term epidemiological studies and informs experimental design and analytical strategies for combining data across cohorts and analytical batches.
Cheung, C.; Glibetic, N.; Maldonado, R.; Bowman, S.; Skaggs, T.; Torres, L.; Perrault Uptmor, K. A.; Weichhaus, M.
Show abstract
BackgroundThe ketogenic diet is being explored as an adjuvant intervention in breast cancer because it lowers circulating glucose and elevates ketone bodies such as {beta}-hydroxybutyrate (BHB), but how individual ER+ breast cancer subtypes adapt to these conditions remains poorly characterized. We examined metabolic responses to BHB supplementation under glucose restriction in two ER+ breast cancer cell lines, asking whether metabolic adaptation patterns differ between models. MethodsMCF-7 and T47D cells were cultured under high glucose, glucose-restricted (5% of standard), or glucose-restricted with 10 mM BHB conditions and profiled by comprehensive two-dimensional gas chromatography-mass spectrometry (GCxGC-MS). Pairwise Welchs t-tests with Benjamini-Hochberg false discovery rate (FDR) correction were applied to identify treatment-responsive metabolites. Targeted assays quantified intracellular glycine, SHMT1 protein, and total branched-chain amino acid (BCAA) concentrations across a BHB dose range (2.5-15 mM). Patient tumor transcriptomic data from TCGA (n=1,084) and paired tumor-normal samples from GSE58135 (n=20) were analyzed for genes involved in one-carbon, ketone body, and BCAA metabolism. ResultsMCF-7 and T47D cells exhibited markedly divergent metabolic responses to BHB. In MCF-7 cells, BHB supplementation produced a broad pattern-level metabolic shift: 75% of detected metabolites trended upward when BHB was added to glucose-restricted cultures (C vs. B comparison), with 1,4-butanediol reaching nominal significance (FC=2.35, p=0.016) and a 4.1-fold trend increase in lactic acid (p=0.11), although no individual metabolite survived FDR correction. T47D cells showed essentially no metabolic response to BHB at the global level. Targeted assays detected an elevation in glycine at 5 mM BHB in both cell lines that did not follow a monotonic dose response and was not accompanied by changes in SHMT1 protein expression. Total BCAA levels were elevated by BHB in T47D cells but remained unchanged in MCF-7 cells. In paired patient samples, OXCT1 (log2FC = -1.41), SHMT1 (log2FC = -1.31), and ACAT1 (log2FC = -1.07) were significantly downregulated in ER+ tumors relative to matched normal tissue (adjusted p < 0.001 for all three). ConclusionsER+ breast cancer cell lines show heterogeneous metabolic responses to BHB supplementation under glucose restriction. The broad pattern of metabolite elevation in MCF-7 but not T47D cells suggests that capacity to utilize ketone bodies as metabolic substrate varies between ER+ models. The downregulation of OXCT1, ACAT1, and SHMT1 in ER+ tumors compared to normal tissue identifies these enzymes as candidate biomarkers that may help stratify which patients are likely to benefit from ketogenic interventions. Findings related to individual metabolites should be regarded as exploratory and require validation in larger, adequately powered cohorts.