Improving Differentiation of Crohns Disease and Ulcerative Colitis Proteomes through Protein-Wide Association Study Feature Selection in Machine Learning
Gorelik, M. G.; Gorelik, A. J.; Fishbein, S. R. S.; Fehlmann, T.; Deepak, P.; Bogdan, R.; Dantas, G.; Jain, U.
Show abstract
Background and AimsDiagnostic differentiation between Crohns disease (CD) and ulcerative colitis (UC) is crucial for timely and suitable therapeutic measures. The current gold standard for differentiating between CD and UC involves endoscopy and histology, which are invasive and costly. We aimed to identify blood plasma proteomic signatures using a Protein-Wide Association Study (PWAS) approach to differentiate CD from UC and evaluate the efficacy of these signatures as features in machine learning (ML) classifiers. MethodsAmong participants (n=1,106; nCD=636; nUC=470) of the Study of a Prospective Adult Research Cohort with IBD (SPARC), plasma protein (n=2,920) levels were estimated using Olink proteomics. A PWAS with Bonferroni correction for multiple testing was used to identify proteins associated with disease states after controlling for age, sex, and disease severity. ML classifiers examined the diagnostic utility of these models. Feature importance was determined via SHapley Additive exPlanations (SHAP) analysis. ResultsThirteen proteins which were significantly differentially abundant in CD vs UC (all |{beta}|s > 0.22, all adjusted p values < 8.42E-06). Random forest models of proteins differentiated between CD and UC with models trained only on PWAS identified proteins (Average ROC-AUC 0.73) outperforming models trained of the full proteome (Average ROC-AUC 0.62). SHAP analysis revealed that Granzyme B, insulin-like peptide 5 (INSL5), and interleukin-12 subunit beta (IL-12B) were the most important features. ConclusionsOur findings demonstrate that PWAS-based feature selection approaches are a powerful method to identify features in complex, noisy datasets. Importantly, we have identified novel peptide based biomarkers such as INSL5, that can be potentially used to complement existing strategies to differentiate between CD and UC.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessment of infant outgrowth of cow's milk allergy in relation to the faecal microbiome and metaproteome 93%
- Comprehensive cell surface proteomics defines markers of classical, intermediate and non-classical monocytes 93%
- Longitudinal Saliva Omics Responses to Immune Perturbation: A Case Study 92%
Similar papers in this journal
- A mass spectrometry-based atlas of extracellular matrix proteins across 25 mouse organs 93%
- Integrated Proteomics analysis of baseline protein expression in pig tissues 92%
- Surveying the vampire bat (Desmodus rotundus) serum proteome: a resource for identifying immunological proteins and detecting pathogens 92%
Similar papers in this journal
- Mass spectrometry-based quantification of proteins and post-translational modifications in dried blood: longitudinal sampling of patients with sepsis in Tanzania 93%
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 93%
- Characterization of cytokine treatment on human pancreatic islets by top-down proteomics 92%
Similar papers in this journal
- Intestinal receptor of SARS-CoV-2 in inflamed IBD tissue is downregulated by HNF4A in ileum and upregulated by interferon regulating factors in colon 92%
- Mapping adaptive immune responses toward fungal antigens in inflammatory bowel disease using T cell repertoire sequencing and phage-immunoprecipitation sequencing 92%
- An extremes of phenotype approach confirms significant genetic heterogeneity in patients with ulcerative colitis. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.