NeTOIF: A Network-based Approach for Time-Series Omics Data Imputation and Forecasting
Shi, M.; Mollah, S.
Show abstract
MotivationHigh-throughput studies of biological systems are rapidly generating a wealth of omics-scale data. Many of these studies are time-series collecting proteomics and genomics data capturing dynamic observations. While time-series omics data are essential to unravel the mechanisms of various diseases, they often include missing (or incomplete) values resulting in data shortage. Data missing and shortage are especially problematic for downstream applications such as omics data integration and computational analyses that need complete and sufficient data representations. Data imputation and forecasting methods have been widely used to mitigate these issues. However, existing imputation and forecasting techniques typically address static omics data representing a single time point and perform forecasting on data with complete values. As a result, these techniques lack the ability to capture the time-ordered nature of data and cannot handle omics data containing missing values at multiple time points. ResultsWe propose a network-based method for time-series omics data imputation and forecasting (NeTOIF) that handle omics data containing missing values at multiple time points. NeTOIF takes advantage of topological relationships (e.g., protein-protein and gene-gene interactions) among omics data samples and incorporates a graph convolutional network to first infer the missing values at different time points. Then, we combine these inferred values with the original omics data to perform time-series imputation and forecasting using a long short-term memory network. Evaluating NeTOIF with a proteomic and a genomic dataset demonstrated a distinct advantage of NeTOIF over existing data imputation and forecasting methods. The average mean square error of NeTOIF improved 11.3% for imputation and 6.4% for forcasting compared to the baseline methods. Contactsmollah@wustl.edu
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Single-cell multi-omics and spatial multi-omics data integration via dual-path graph attention auto-encoder 96%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 94%
- Gene Expression Prediction from Histology Images via Hypergraph Neural Networks 94%
Similar papers in this journal
- DAGBagM: Learning directed acyclic graphs of mixed variables with an application to identify prognostic protein biomarkers in ovarian cancer 95%
- Single-Cell Classification Using Graph Convolutional Networks 95%
- DeepMF: Deciphering the Latent Patterns in Omics Profiles with a Deep Learning Method 95%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
- Minn: A Metabolic-Informed Neural Network For Integrating Omics Data Into Genome-Scale Metabolic Modeling 93%
- Wide and Deep Learning for Automatic Cell Type Identification 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.