Web Browsing Using Machine Learning on Text Data

Dunja Mladenić · Studies in fuzziness and soft computing · 2003

Web browsing is gaining popularity with the growing number of Web users, especially for a casual usage of the Web, when the user does not have a precise query in mind. By observing the user’s behavior when browsing, we build a model of promising hyperlinks and use it to highlight hyperlinks on the requested Web pages. In order to do that, we propose text-learning methods for handling high dimensional problems (having several tens of thousands of features) with highly unbalanced class distribution (more than 90% of examples having the majority class value). The reported experimental results on the user modeling problem are consistent with the extensive experimental results that were performed on a related problem of modeling Web document content category by using hyperlink to the document. The results show that when modeling by Naive Bayesian classifier, it is highly important how we select the features to be used in the model. Namely, the best performing feature selection in our experiments on Personal WebWatcher data is when the features are scored according to Odds Ratio and only a small number of the best features is used for learning.

Read the paper · More papers on PaperTik