Back

A Proposal of New Feature Selection Method Sensitive to Outliers and Correlation

Demirarslan, M.; Suner, A.

2021-03-12 bioinformatics
10.1101/2021.03.11.434934 bioRxiv
Show abstract

In disease diagnosis classification, ensemble learning algorithms enable strong and successful models by training more than one learning function simultaneously. This study aimed to eliminate the irrelevant variable problem with the proposed new feature selection method and compare the ensemble learning algorithms classification performances after eliminating the problems such as missing observation, classroom noise, and class imbalance that may occur in the disease diagnosis data. According to the findings obtained; In the preprocessed data, it was seen that the classification performance of the algorithms was higher than the raw version of the data. When the algorithms classification performances for the new proposed advanced t-Score and the old t-Score method were compared, the feature selection made with the proposed method showed statistically higher performance in all data sets and all algorithms compared to the old t-Score method (p = 0.0001).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.