Hybrid Variable Selection Approach to Analyse High Dimensional Dataset
Siva Subramanian R, K. Sudha, Venkata Ramana K, S. SivaKumar, R. Nithyanandhan, M. Nalini · 2023
In light of the heterogeneous mode of collection and high dimensionality, most real-time datasets collected may contain irrelevant, correlated, noisy, and missing variables. When these high-dimensional datasets are used with machine learning, the resulting prediction about the dataset is inaccurate. A technique called feature selection is being usedin an effort to solve the above problem. FS examines the relevant variable from the entire dataset and selects the variable based on the evaluation strategy used. In nature, there are three types of FS techniques. The first one is the first approach, which is fast, but not effective. Likewise, the wrapper approach is an efficient one but the problem is taking a high computation time. This line of research aims to develop a new kind of FS known as hybrid variable selection in order to get around the drawbacks that are associated with the filter and wrapper approach. The purpose of HVS is to extract from the full dataset an efficient and relevant variable subset using the advantages offered by the filter and wrapper approaches. An empirical method is implemented utilizing the UCI-collected consumer dataset. The proposed methodology is implemented by using three different perspectives: first, the filter technique, next, the wrapper strategy, and finally, a hybrid approach. Further, variable subset captured from the three different approaches is modelled with the ML classifier separately. Here Naive Bayes, ML is applied. A variety of validity measures are used to forecast and compare experimental results. Final analysis results suggest the suggested HVS performs efficiently in comparison to the filter and wrapper technique.