Simple Vector Representations of E-Commerce Products

Abe Gracia Vallerian Tenno Siswanto, Lilian Tjong, Yordan Saputra · 2018

Product similarity is a fundamental function in many aspects of e-commerce. To calculate the similarity, first, we need to build a representation that draws similar products closer in a vector space, i.e. embedding. In our experiments, we create product embeddings of n-dimension using unstructured texts like product names as the dataset. After constructing the word embeddings, using Word2Vec, we combine the vectors with two different methods: 1) unweighted average and 2) weighted average with noise removal. For evaluation, we use two different tasks, which are calculating product similarity and integrating the embedding as a feature for product category prediction. The result shows that the second method is better at doing simple task of calculating product similarity with more than 92 % accuracy. As for product category prediction, the first method consistently gives the best result with around 71 % accuracy and 86 % recall.

Read the paper · More papers on PaperTik