Multi-view confounder detection for biomedical studies
Wallner, M. A.; Kersten, N.; Pfeifer, N.
Show abstract
In many biomedical studies an important first step is checking for confounding factors. For association studies, confounding can for example be caused by ethnic differences in the case and control groups. In many other settings there might be confounding factors like batch effects or founder effects that also need to be detected and controlled for1. Detecting confounding for data from one data source is well established (e.g., genomics data). Since more and more studies are now based on data from multiple data modalities (e.g., multi-omics), we evaluated whether multi-view confounder detection can benefit from state-of-the-art methods for multi-view data integration. Especially for clustering of multi-omics data, it has been shown that these methods can perform better than methods that treat the data modalities separately2. Our results show that multi-view confounder analysis is possible and that building on multi-view data integration methods is better than treating the different data modalities separately.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Optimal Transport improves cell-cell similarity inference in single-cell omics data 94%
- Per-sample standardization and asymmetric winsorization lead to accurate clustering of RNA-seq expression profiles 93%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 93%
Similar papers in this journal
- MarcoPolo: a clustering-free approach to the exploration of differentially expressed genes along with group information in single-cell RNA-seq data 93%
- Integration of Gene Expression and DNA Methylation Data Across Different Experiments 93%
- Assessing the impact of transcriptomics data analysis pipelines on downstream functional enrichment results 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.