WebRank: Language-Independent Extraction of Keywords from Webpages

Himat Ali Shah, Radu Mariescu-Istodor, Pasi Fränti · 2021

We present a supervised method for keyword extraction from webpages. The method divides the HTML page into meaningful segments using document object model (DOM) and calculates a language independent feature vector for each word. Based on these, we generate a classification model that gives a likelihood for a word to be a keyword. The most likely words are then selected. We analyze the usefulness of the features on different datasets (news articles and service web pages) and compare different classification methods for the task. Results show that random forest performs best and provides up to 27.8 %- unit improvement compared to the best existing method.

Read the paper · More papers on PaperTik