The results of Transcriptome-wide Mendelian Randomization (TWMR) in large-scale populations can directly validate, across scales, the results of causal inference from deep learning combined with double machine learning on single-cell transcriptomes of human samples.
ye, w.; Jiang, X.; Shen, F.
Show abstract
ObjectiveAiming at the core problems prevalent in biomedical research, including the "translational distance", the difficulty in aligning cross-scale studies, and the lack of direct validation of single-cell systems biology models in human samples, this study aims to verify whether the results of transcriptome-wide Mendelian randomization (TWMR) based on large-scale populations are consistent with the causal inference results of deep learning combined with double machine learning (DML) using single-cell transcriptome data from human samples, to clarify whether statistical biology and systems biology can converge to the same biological truth, and provide methodological support for mechanism dissection and precision medicine research of complex diseases such as rheumatoid arthritis (RA). MethodsThis study integrated multi-omics data to conduct a two-stage causal inference and cross-scale validation analysis. In the first stage, based on the summary statistics of RA genome-wide association study (GWAS) from 456,348 individuals of European ancestry in the UK Biobank (UKB), and cis-expression quantitative trait locus (cis-eQTL) data from 31,684 individuals in the eQTLGen Consortium, a two-sample Mendelian randomization approach was adopted. Transcriptome-wide causal effect analysis was performed using the inverse-variance weighted (IVW) method, MR Egger regression, and weighted median method, and gene-level causal effect values were obtained after strict quality control and multiple testing correction. In the second stage, based on single-cell RNA sequencing (scRNA-seq) data from RA patients and healthy controls (RA group: 11 samples, 211,867 cells; Healthy control group: 38 samples, 456,631 cells), after preprocessing via the Seurat pipeline, batch effect correction, and cell type annotation, a hierarchical deep neural network was constructed to complete feature compression of high-dimensional expression data, and the DML framework was used to estimate the causal effects of genes on RA disease status. Finally, Pearson correlation analysis was performed to conduct cell type-specific cross-scale validation of gene-level causal effect values obtained by the two methods, and the validated model was used to quantify the causal effects of 16 RA-related pathways from the Reactome database. ResultsThis study confirmed that the gene causal effect values obtained from large-scale population TWMR analysis were significantly correlated with those calculated by the deep learning combined with DML model based on single-cell transcriptome data. Among them, the correlation was extremely significant (p<0.001) in core naive B cells (r=0.202, p=3.2e-05, n=414) and core naive CD4 T cells (r=0.102, p=0.037, n=412). The validated DML model successfully quantified the cell type-specific causal effect values of 16 RA-related signaling pathways. ConclusionStatistical biology and systems biology can converge to the same biological truth. The cross-scale consistency between the two can significantly shorten the "translational distance" in biomedical research, and realizes the direct validation of the single-cell systems biology causal model of human samples based on large-scale population genetic data, getting rid of the excessive dependence on animal/cell experimental models in traditional research. This research paradigm not only provides a new path for mechanism dissection and therapeutic target screening of complex diseases such as RA, but also provides a feasible solution for rare disease research to break through the limitation of GWAS sample size, and lays an important theoretical and methodological foundation for constructing standardized systems biology models of human complex diseases and promoting the development of precision medicine.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 94%
- Noise regularization removes correlation artifacts in single-cell RNA-seq data preprocessing 94%
- DAISM-DNNXMBD: Highly accurate cell type proportion estimation with in silico data augmentation and deep neural networks 93%
Similar papers in this journal
- Single cell sequencing analysis uncovers genetics-influenced CD16+monocytes and memory CD8+T cells involved in severe COVID-19 94%
- scGRNom: a computational pipeline of integrative multi-omics analyses for predicting cell-type disease genes and regulatory networks 94%
- Discovery of CD80 and CD86 as recent activation markers on regulatory T cells by protein-RNA single-cell analysis 94%
Similar papers in this journal
- Unravelling the shared genetic mechanisms underlying 18 autoimmune diseases using a systems approach 94%
- Cross-Tissue Transcriptomic Analysis Leveraging Machine Learning Approaches Identifies New Biomarkers for Rheumatoid Arthritis 94%
- BCR, not TCR, repertoire diversity is associated with favorable COVID-19 prognosis 94%
Similar papers in this journal
- Knowledge-primed neural networks enable biologically interpretable deep learning on single-cell sequencing data 94%
- Functional enrichment of alternative splicing events with NEASE reveals insights into tissue identity and diseases 94%
- A time-resolved meta-analysis of consensus gene expression profiles during human T-cell activation 94%
Similar papers in this journal
- Using Genetics, Genomics, and Transcriptomics to Identify Therapeutic Targets in Juvenile Idiopathic Arthritis 94%
- Genetic analyses of inflammatory polyneuropathy and chronic inflammatory demyelinating polyradiculoneuropathy identified candidate genes 94%
- Discovery of disease-associated cellular states using ResidPCA in single-cell RNA and ATAC sequencing data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.