Back

Hybrid TF-IDF and SBERT Feature Engineering with SMOTE-Tomek Resampling for Enhanced Software Requirements Classification

Ali, H.

2026-01-14 bioinformatics
10.64898/2026.01.13.699413 bioRxiv
Show abstract

Software requirements classification remains one of the important challenges in requirements engineering. Engineering that affects the smoothness of project success about software development life cycles. in this paper, a novel hybrid solution is being presented that beats the benchmarks set by previous approaches using TF-IDF. The vectorization, coupled with SBERT embeddings, has been further reinforced by application of SMOTE-Tomek resampling. The solution addresses typical limitations detected in Ors baseline solution, which achieved 76.16% {+/-} 2.58% accuracy of the model using the conventional SMOTE-Tomek pre-processing with logistic regression. Via comprehensive experimentation on the PROMISE dataset containing 969 categorized requirements accross 12 classes, our hybrid feature engineering approach has shown significant improvements: Random Forest achieves. Thus, the character CNN records an accuracy of 74.40% {+/-} 3.00% versus the baseline 60.37% {+/-} 2.76%, while critically, Naive Bayes shows remarkable improvement from complete failure 0.00% to 53.97% {+/-} 5.08% accuracy. The integration of semantic contextual understanding through SBBERT, with statistical term importances via TF-IDF, forms a robust representation space that captures the synthetic pattern along with semantic relations in the texts of requirements. Our methodology uses stratified 10-fold cross-validation with Matthews Correlation Coefficient (MCC). Evaluation, Ensuring reliable performance assessment across imbalanced classes. The results demonstrate that hybrid feature engineering, when combined with appropriate resampling techniques, provides a more effective solution for automated software requirement classification than traditional pre-processing methods alone.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.