Back

A Robust and Interpretable Feature Engineering Approach for Low-Data Biological Classification

Gora, S.; Dervishi, E.

2026-01-02 bioinformatics
10.64898/2026.01.02.697379 bioRxiv
Show abstract

Accurate biological classification often faces challenges from high-dimensional data and limited samples, leading to model overfitting and poor interpretability. This study introduces Directional Flow Embedding (DFE), a novel feature engineering and embedding method designed to overcome these issues. DFE transforms raw biological data into three concise and biologically interpretable features: Directional Flow Score (DFS), Heterogeneity Index (HI), and Local Density Estimate (LDE). It achieves this by robustly determining a global direction vector from class means, enabling the projection of samples onto this principal axis, quantifying their deviation, and incorporating local density information. Evaluated on real-world TCGA-BRCA and TCGA-LUAD RNA-sequencing datasets, DFE consistently demonstrated superior performance. It significantly outperformed traditional methods and strong non-linear models. Ablation studies confirmed the synergistic contribution of all three DFE features, while sensitivity analysis revealed its robustness. DFEs inherent interpretability, strong generalizability, and computational efficiency make it a valuable tool for robust and transparent biological classification, thereby advancing critical applications in biomedical research and clinical decision-making.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.