Revisiting Logistic Regression for High-Dimensional Gene Expression Data
Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.
Show abstract
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 94%
- Estimating Microbial Interaction Network:Zero-inflated Latent Ising Model Based Approach 90%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 89%
Similar papers in this journal
- Projection in genomic analysis: A theoretical basis to rationalize tensor decomposition and principal component analysis as feature selection tools 94%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 94%
- Identification of high-risk COVID-19 patients using machine learning 94%
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 94%
- Machine Learning Interpretability Methods to Characterize the Importance of Hematologic Biomarkers in Prognosticating Patients with Suspected Infection 94%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 94%
Similar papers in this journal
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 95%
- Molecular Group and Correlation Guided Structural Learning for Multi-Phenotype Prediction 94%
- High Dimensionality Reduction by Matrix Factorization for Systems Pharmacology 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.