A data-fusion approach to identifying developmental dyslexia from multi-omics datasets
Carrion, J. T.; Nandakumar, R.; Shi, X.; Gu, H.; Kim, Y.; Raskind, W. H.; Peter, B.; Dinu, V.
Show abstract
This exploratory study tested and validated the use of data fusion and machine learning techniques to probe high-throughput omics and clinical data with a goal of exploring the etiology of developmental dyslexia. Developmental dyslexia is the leading learning disability in school aged children affecting roughly 5-10% of the US population. The complex biological and neurological phenotype of this life altering disability complicates its diagnosis. Phenome, exome, and metabolome data was collected allowing us to fully explore this system from a behavioral, cellular, and molecular point of view. This study provides a proof of concept showing that data fusion and ensemble learning techniques can outperform traditional machine learning techniques when provided small and complex multi-omics and clinical datasets. Heterogenous stacking classifiers consisting of single-omic experts/models achieved an accuracy of 86%, F1 score of 0.89, and AUC value of 0.83. Ensemble methods also provided a ranked list of important features that suggests exome single nucleotide polymorphisms found in the thalamus and cerebellum could be potential biomarkers for developmental dyslexia and heavily influenced the classification of DD within our machine learning models.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Applying machine learning in motor activity time series of depressed bipolar and unipolar patients. 92%
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 92%
- MultipleTesting.com: a tool for life science researchers for multiple hypothesis testing correction 92%
Similar papers in this journal
- Novel ratio-metric features enable the identification of new driver genes across cancer types 92%
- Optimised multiplex amplicon sequencing for mutation identification using the MinION nanopore sequencer 91%
- Effectiveness of three bioinformatics tools in the detection of ASD candidate variants from whole exome sequencing data 91%
Similar papers in this journal
Similar papers in this journal
- Protein profiling of WERI RB1 and etoposide resistant WERI ETOR reveals new insights into topoisomerase inhibitor resistance in retinoblastoma 92%
- Multi-run Concrete Autoencoder to Identify Prognostic lncRNAs for 12 Cancers 91%
- Identification of ATP2B4 regulatory element containing functional genetic variants associated with severe malaria 90%
Similar papers in this journal
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 92%
- FARFOOD: A database of potential interactions between food compounds and drugs. 89%
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.