Frisky CALF outruns LASSO
Jeffries, C. D.; Ford, J. R.; Tilson, J. L.; Perkins, D. O.; Bost, D.; Filer, D.; Wilhelmsen, K. C.
Show abstract
Regularized regression analysis is a mature analytic approach to identify weighted sums of variables that predict outcomes. Typically, the number of subjects (N) is smaller than the number of predictors (p). Here, we present a novel coarse approximation linear function (CALF) to frugally select important predictors, build linear models, and discover causality. CALF is a linear regression strategy applied to normalized data that employs only a few (such as 2 to 20) nonzero weights, each +1 or -1. Metrics can be Welch t-test p-value, area under curve (AUC) of receiver operating characteristic, or Pearson correlation, depending upon data type and user preferences. For quantitative approximations, a linear fit (adding an intercept value and rescaling {+/-}1 weights by a common multiplier) can be added to optimize mean squared error (MSE). Real medical data of five types were used to generate examples with goals of binary classification or real variable approximation. Predictors considered were real, sets of integers, or ternary values of single nucleotide polymorphisms. When applied to real data, CALF approximations outperformed in p-value, AUC, correlation, or MSE a popular regularized linear regression algorithm, namely, basic LASSO. It appears that using LASSO without considering CALF might risk wasting resources. AvailabilityR version: Comprehensive R Archive (CRAN): https://cran.r-project.org/web/packages/CALF/index.html Python 3.x version: GitHub: https://github.com/jorufo/CALF_Python Contactclark_jeffries@med.unc.edu
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Soft Windowing Application to Improve Analysis of High-throughput Phenotyping Data 95%
- dsMTL - a computational framework for privacy-preserving, distributed multi-task machine learning 94%
- Nearest-neighbor Projected-Distance Regression (NPDR) for detecting network interactions with adjustments for multiple tests and confounding 94%
Similar papers in this journal
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 94%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 93%
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 95%
- Topic Modeling analysis of the Allen Human Brain Atlas 94%
- Accurate Prediction of Breast Cancer Survival through Coherent Voting Networks with Gene Expression Profiling 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.