CrosSplice: A Pipeline for Identifying Rare Splice-Site Creating Variants from Cross-Tissue Transcriptome Data
Yano, Y.; Okada, A.; Ono, M.; Mateos, R. N.; Nakayama, T.; Shiraishi, Y.
Show abstract
BackgroundDespite their profound impact on patients lives, most rare and intractable diseases still lack established treatments. Genomic variants that disrupt normal splicing by creating novel splice sites (splice-site creating variants, SSCVs) substantially contribute to the pathogenesis of those conditions. Deep intronic SSCVs are particularly amenable to antisense oligonucleotide (ASO)-mediated splice modulation, yet many of them remain undetected by conventional genomic analyses. Existing approaches to identify SSCVs, including splicing quantitative trait loci (sQTL) analyses and machine learning-based methods, each show limitations in sensitivity or accuracy, hindering the comprehensive identification of clinically actionable SSCVs. MethodsWe developed CrosSplice, a novel pipeline that integrates machine learning-based predictions with statistical association testing to robustly identify SSCVs. By leveraging cross-tissue transcriptome data and aggregating splicing signals across multiple tissues, CrosSplice enables comprehensive detection of SSCVs, including rare and tissue-specific variants often missed by conventional methods. ResultsApplying CrosSplice to the GTEx dataset (8,656 transcriptomes of 54 tissues from 479 postmortem donors), we identified 1,743 significant SSCVs, 65% of which were deep intronic. Among these, 185 SSCVs were listed in ClinVar, including five pathogenic or likely pathogenic variants. CrosSplice also discovered a novel deep intronic SSCV in PLA2G6, the gene responsible for infantile neuroaxonal dystrophy. We experimentally confirmed that ASOs successfully corrected the aberrant splicing pattern induced by this variant. ConclusionCrosSplice substantially extends the detectable landscape of SSCVs by capturing rare and tissue-specific variants, uncovering pathogenic and therapeutically actionable SSCVs that are frequently overlooked by existing methods. The resulting SSCV catalogue provides a platform for systematic discovery of ASO targets and advances opportunities for precision therapies in rare and intractable diseases.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MRSD: a novel quantitative approach for assessing suitability of RNA-seq in the clinical investigation of mis-splicing in Mendelian disease 96%
- Transcriptome-wide outlier approach identifies individuals with minor spliceopathies 95%
- High-throughput splicing assays identify missense and silent splice-disruptive POU1F1 variants underlying pituitary hormone deficiency 95%
Similar papers in this journal
Similar papers in this journal
- Systematic analysis of genetic and phenotypic characteristics reveals antisense oligonucleotide therapy potential for one-third of neurodevelopmental disorders 97%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 95%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 93%
Similar papers in this journal
- Alternative splicing is coupled to gene expression in a subset of variably expressed genes 93%
- WEGS: a cost-effective sequencing method for genetic studies combining high-depth whole exome and low-depth whole genome 92%
- Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort 92%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- Evaluation of imputation performance of multiple reference panels in a Pakistani population 93%
- Identification and validation of novel candidate risk genes in endocytic vesicular trafficking associated with esophageal atresia and tracheoesophageal fistulas 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.