A novel method to guide biomarker combinations to optimize the sensitivity
Irajizad, E.; Long, J. P.; Do, K.-A.; Fahrmann, J.; Hanash, S.; Ghasemi, S. M.
Show abstract
Logistic regression has demonstrated its utility in classifying binary labeled datasets through the maximum likelihood approach. However, in numerous biological and clinical contexts, the aim is often to determine coefficients that yield the highest sensitivity at the pre-specified specificity or vice versa. Therefore, the application of logistic regression is limited in such settings. To this end, we have developed an improved regression framework, SMAGS, for binary classification that, for a given specificity, finds the linear decision rule that yields the maximum sensitivity. Furthermore, we employed the method for feature selection to find the features that are satisfying the sensitivity maximization goal. We compared our method with normal logistic regression by applying it to real clinical data as well as synthetic data. In the real application data (colorectal cancer dataset), we found 14% improvement of sensitivity at 98.5% specificity. Availability and implementationSoftware is made available in Python (https://github.com/smahmoodghasemi/SMAGS)
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 96%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 96%
- A machine-learning Approach for Stress Detection Using Wearable Sensors in Free-living Environments 95%
Similar papers in this journal
Similar papers in this journal
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 95%
- Estimating Microbial Interaction Network:Zero-inflated Latent Ising Model Based Approach 93%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.