Back

Revisiting Logistic Regression for High-Dimensional Gene Expression Data

Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.

2026-07-24 bioinformatics
10.64898/2026.07.20.739668 bioRxiv
Show abstract

Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.