A machine-learning evaluation of biomarkers designed for the future of precision medicine
Climer, S.
Show abstract
Precision medicine is cognizant of the impact of genetics and environments on subtypes of heterogeneous diseases and aims to identify, diagnose, and treat each subtype appropriately. Real-valued biomarkers, such as protein levels in plasma, are key for practical subtype diagnoses and hold potential to elucidate subtypes and illuminate promising drug targets. Biomarkers that are common across all subtypes have been discovered using fold change (FC) and the area under the receiver operating characteristic curve (AUC). However, FC and AUC fail to identify biomarkers for subtypes when they comprise less than half of the disease group. We present here a machine-learning biomarker evaluation method based on clustering of the data points, referred to as Difference in Bicluster Distances (DBD). We contribute efficient, yet optimal, software coupled with rigorous validation techniques, and demonstrate our approach on a late-onset Alzheimer disease (AD) gene expression dataset. Our trials produced four significant genes and appropriate thresholds for biomarker diagnostics. While none of these genes were identified as significant by either FC or AUC for the given dataset, the genes have been independently associated with AD or neurological disorders by other groups using completely independent means. In summary, DBD provides a unique and effective method for screening real-valued data to identify biomarkers associated with subtypes of heterogeneous diseases.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 93%
- Co-localized SNPs Affecting the Expression of Taste Perception Genes are linked to Alzheimer's Disease 91%
- Physics-informed machine learning for automatic model reduction in chemical reaction networks 91%
Similar papers in this journal
- CharMark: A Markov Approach to Linguistic Biomarkers in Dementia 90%
- Using Machine Learning to Predict Mortality for COVID-19 Patients on Day Zero in the ICU 89%
- Retinal Vascular Measures from Diabetes Retinal Screening Photographs and Risk of Incident Dementia in Type 2 Diabetes: A GoDARTS Study 88%
Similar papers in this journal
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 89%
- Modeling physician variability to prioritize relevant medical record information 89%
Similar papers in this journal
- Identifying and ranking potential driver genes of Alzheimer's Disease using multi-view evidence aggregation 92%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 91%
- PathWalks: Identifying pathway communities using a disease-related map of integrated information 91%
Similar papers in this journal
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 94%
- Identification of functionally connected multi-omic biomarkers for Alzheimer’s Disease using modularity-constrained Lasso 93%
- Examining heterogeneity in dementia using data-driven unsupervised clustering of cognitive profiles 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.