Keyword Extraction in News Articles Using Table based K Nearest Neighbors
Duke Taeho Jo · 2018
In this research, we propose the KNN version where words are encoded into tables, instead of numerical vectors, as the approach to the keyword extraction. The keyword extraction is mapped into a binary classification task within a domain, and the task should be distinguished from the topic based word categorization. In this research, words are encoded into tables each of which consists of entries of text identifiers and their weights, the KNN algorithm is modified by adopting the proposed similarity metric, and it is applied to the keyword extraction which is mapped into a binary classification. It is validated empirically that the proposed KNN version is better than the traditional version in extracting keywords from a text which is tagged with its own domain. In future, we will connect the task with the text categorization, in order to process texts which are untagged with their domains.