Proportioning documents over categories based on word embeddings
Jiahua Du, Jing He · Proceedings of the Australasian Computer Science Week Multiconference · 2016
News articles, even in the same category, often interweave with multiple stories and topics. In fact, exploring semantic components of textual documents is challenging, and assigning a document to one single category may appear inadequate. This paper employs word embeddings techniques to proportion documents over categories via semantic vector operations. Word embeddings are useful in capturing word semantics. The proposed method first prepares keyword collections for each category by identifying representative words from training set; calculates the semantic similarity between each new document and the word collections of each category; and analyzes the category distribution of the document using percentages. In the distribution, larger values indicate higher similarities between the category and document. The proposed method is evaluated by comparing categories yielding top values in the distribution of a document with its original category. Experimental results show that the proposed method can efficiently analyze the semantic components of news articles.