Comprehensive Analysis of Multi-Omics Vaccine Response Data Using MOFA and Stabl Algorithms
Gupta, A.; Abe, K.; Maecker, H.
Show abstract
FluPRINT is a multi-omics dataset that measures donors protein expression and cell counts across various assays. Donors were also assigned a binary value (0 or 1), being labeled as high responders (1) if they had a fold change [≥] 4 of the antibody titer for hemagglutinin inhibition (HAI) from day 0 to day 28, and low responders otherwise (0). In this project, we used the MOFA and Stabl algorithms to analyze FluPRINT, estimate the population structure from the data, and identify the most important features for predicting response to the vaccine. The preprocessing of the dataset included removing repeat features, scaling by assay, and removing outliers. Since Stabl does not directly address missing values, features with high amounts of missing values were removed and the remaining were ignored. MOFA identified the top feature in structure extraction as IL neg 2 CD4 pos CD45RA neg pSTAT5. MOFA explains well the variance of the data while also choosing features that have good significance, as illustrated by their significant p-values (p < 0.05). Stabl found the top feature for explaining the outcome to be CD33- CD3+ CD4+ CD25hiCD127low CD161+ CD45RA+ Tregs, which matched the top result of previously published analysis. MOFAs features achieved an AUROC of 0.616 (95% CI of 0.426-0.806), and Stabls achieved an AUROC of 0.634 (95% CI of 0.432-0.823). Our research addresses a key knowledge gap: understanding how these fundamentally different analytical approaches perform when analyzing the same complex dataset. Our exploration evaluates their respective strengths, limitations, and biological insights and provides guidance on using MOFA and Stabl to find the best predictive cell subsets and features for understanding large immunological multi-omics data. The code for this project can be found at https://github.com/aanya21gupta/fluprint.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Immune-Based Prediction of COVID-19 Severity and Chronicity Decoded Using Machine Learning 95%
- MotifBoost: k-mer based data-efficient immune repertoire classification method 93%
- Development and Validation of Multivariable Prediction Models of Serological Response to SARS-CoV-2 Vaccination in Kidney Transplant Recipients 93%
Similar papers in this journal
- FaDA: A Shiny web application to accelerate common lab data analyses 95%
- Imbalanced Machine Learning Classification Models For Removal Biosimilar Drugs And Increased Activity In Patients With Rheumatic Diseases 93%
- A distinct four-value blood signature of pyrexia under combination therapy of malignant melanoma with BRAF/MEK-inhibitors evidenced by an algorithm-defined pyrexia score 93%
Similar papers in this journal
- A hybrid method for discovering interferon-gamma inducing peptides in human and mouse 93%
- Distinct SARS-CoV-2 Antibody Reactivity Patterns in Coronavirus Convalescent Plasma Revealed by a Coronavirus Antigen Microarray 93%
- In silico tool for Predicting, Designing and Scanning IL-2 inducing peptides 93%
Similar papers in this journal
- Enhanced Annotation of CD45RA to Distinguish T cell Subsets in Single Cell RNA-seq via Machine Learning 96%
- LMPred: Predicting Antimicrobial Peptides Using Pre-Trained Language Models and Deep Learning 93%
- HLA-EpiCheck: A B-cell epitope prediction tool for HLA proteins using molecular dynamics simulation data. 93%
Similar papers in this journal
- Classification of human white blood cells using machine learning for stain-free imaging flow cytometry 95%
- Integration, exploration, and analysis of high-dimensional single-cell cytometry data using Spectre 92%
- A 30-Color Full-Spectrum Flow Cytometry Panel to Characterize the Immune Cell Landscape in Spleen and Tumors within a Syngeneic MC-38 Murine Colon Carcinoma Model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.