A systematic review on machine learning approaches in the diagnosis of rare genetic diseases
Roman-Naranjo, P.; Parra-Perez, A. M.; Lopez-Escamez, J. A.
Show abstract
BackgroundThe diagnosis of rare genetic diseases is often challenging due to the complexity of the genetic underpinnings of these conditions and the limited availability of diagnostic tools. Machine learning (ML) algorithms have the potential to improve the accuracy and speed of diagnosis by analyzing large amounts of genomic data and identifying complex multiallelic patterns that may be associated with specific diseases. In this systematic review, we aimed to identify the methodological trends and the ML application areas in rare genetic diseases. MethodsWe performed a systematic review of the literature following the PRISMA guidelines to search studies that used ML approaches to enhance the diagnosis of rare genetic diseases. Studies that used DNA-based sequencing data and a variety of ML algorithms were included, summarized, and analyzed using bibliometric methods, visualization tools, and a feature co-occurrence analysis. FindingsOur search identified 22 studies that met the inclusion criteria. We found that exome sequencing was the most frequently used sequencing technology (59%), and rare neoplastic diseases were the most prevalent disease scenario (59%). In rare neoplasms, the most frequent applications of ML models were the differential diagnosis or stratification of patients (38.5%) and the identification of somatic mutations (30.8%). In other rare diseases, the most frequent goals were the prioritization of rare variants or genes (55.5%) and the identification of biallelic or digenic inheritance (33.3%). The most employed method was the random forest algorithm (54.5%). In addition, the features of the datasets needed for training these algorithms were distinctive depending on the goal pursued, including the mutational load in each gene for the differential diagnosis of patients, or the combination of genotype features and sequence-derived features (such as GC-content) for the identification of somatic mutations. ConclusionsML algorithms based on sequencing data are mainly used for the diagnosis of rare neoplastic diseases, with random forest being the most common approach. We identified key features in the datasets used for training these ML models according to the objective pursued. These features can support the development of future ML models in the diagnosis of rare genetic diseases.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High Precision Characterization Of Rccx Rearrangements In A 21-Hydroxylase Deficiency Latin American Cohort Using Oxford Nanopore Long Read Sequencing 92%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 92%
- Altered gene expression profiles impair the nervous system development in individuals with 15q13.3 microdeletion 92%
Similar papers in this journal
- Characterizing sensitivity and coverage of clinical WGS as a diagnostic test for genetic disorders 94%
- GeneTerpret: a customizable multilayer approach to genomic variant prioritization and interpretation 93%
- Genome-wide survey of tandem repeats by nanopore sequencing shows that disease-associated repeats are more polymorphic in the general population 89%
Similar papers in this journal
- Defining the Critical Components of Informed Consent for Genetic Testing: A Delphi Study 90%
- Identification of Somatic Structural Variants in Solid Tumors By Optical Genome Mapping 89%
- Comprehensive profiling of genomic and transcriptomic differences between risk groups of lung adenocarcinoma and lung squamous cell carcinoma 89%
Similar papers in this journal
- Genetic Insight into Birt-Hogg-Dubé syndrome in Indian patients reveals novel mutations in FLCN 92%
- DDIEM: Drug Database for Inborn Errors of Metabolism 91%
- The COVID-19 pandemic impact on continuity of care provision on rare brain diseases and on Ataxia, Dystonia and PKU. A scoping review protocol 89%
Similar papers in this journal
- Accelerate the discovery of genetic variants in mitochondrial diseases with VIOLA: Variant PrIOritization using Latent space 93%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 90%
- In silico tools for accurate HLA and KIR inference from clinical sequencing data empower immunogenetics on individual-patient and population scales 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.