A method for automatic extraction of multiword units representing business aspects from user reviews
Olga Vechtomova · Journal of the Association for Information Science and Technology · 2014
The article describes a semi‐supervised approach to extracting multiword aspects of user‐written reviews that belong to a given category. The method starts with a small set of seed words, representing the target category, and calculates distributional similarity between the candidate and seed words. We compare 3 distributional similarity measures (Lin's,Weeds's, and balAPinc), and a document retrieval function,BM25, adapted as a word similarity measure. We then introduce a method for identifying multiword aspects by using a combination of syntactic rules and a co‐occurrence association measure. Finally, we describe a method for ranking multiword aspects by the likelihood of belonging to the target aspect category. The task used for evaluation is extraction of restaurant dish names from a corpus of restaurant reviews.