An Integrated Framework for Preprocessing and Clustering Strategies in WEB Mining

S. Ganeshmoorthy · 2024

Numerous organizations position a premium through the fast expansion of web-based services, web-based IT infrastructures, and online commerce facilitating the web "Log-File (LF)" evaluations. Web LF encompasses data about the user's online activity. The most commonly employed LF structure is the standard LF format, which is generated by web servers to record all the HTTP queries made to a website. There are certain ambiguous and uncertain information found in LFs that could interfere with the "Web-Usage Mining (WUM)" process. Varied preprocessing procedures are applied to LFs from diverse sources to provide comprehensive "DataSet (DS)". This aspect contributes to the poor accuracy along with the prolonged time with the present methods. The suggested research designed an integrated framework that uses "Hashing-based Damaged Detection (HDD)" to obtain preprocessing, "Hashing-based Dimensionality Reduction (HDR)" to perform data filtering, and "Network Based Aggregative Clustering (NAC)" over users and sessions classification to fix those problems and expand user needs' being available. With the implementation of NASA's Web LF "DataBase (DB)", the suggested (HDD-HDR-NAC) system prioritizes the user's LF to improve the integrity of data. Organizations or website providers could benefit in an innovative way from the suggested (HDD-HDR-NAC) framework, which groups users according to similarity measures, in terms of the amount of time and information contributed during each session. Performance indicators including "Text-Categorization," "Noise-Classification Ratio," and "Accuracy" reveal that the suggested (HDD-HDR-NAC) framework outperforms the current OKNN-FF framework.

Read the paper · More papers on PaperTik