Constraining Weighted Word Co-occurrence Frequencies in Word Embeddings
Paula Lauren · 2021 IEEE International Conference on Big Data (Big Data) · 2021
Weighted word co-occurrence frequencies are considered the bedrock of word embeddings. Also known as a low-dimensional numerical representation, word embeddings capture word pair frequencies extracted from a corpus in an unsupervised manner. The rendering of word embeddings can be considered a two-step process with the first s tep i nvolving t he building of the word context matrix then using a matrix factorization method to reduce the dimensionality. In this research study, word embeddings are constructed from scratch in building the word context matrix and Truncated Singular Value Decomposition is applied to the matrix. Five experimental values are defined for constraining the frequency weights in the word embeddings, which are then evaluated in word similarity and sequence labeling tasks with results reported. The word similarity task shows comparable results across all experimental constraint values. Overall comparable results are also achieved in the sequence labeling task. The experiments conducted in this study have shown promising results, which will entail future work with evaluation on other tasks.