Hierarchical Paragraph Vectors

Lukas Elmer · 2015

Most standard machine learning algorithms require fixed-length, low-dimensional vectors to perform well. However, when working with text, such representations are difficult to obtain. In 2014, Le and Mikolov presented a novel method for generating so-called word embeddings [16]. Their work represents the foundation for this master thesis. The main goal of this thesis is to extend and generalize the word em-bedding model to a hierarchical paragraph vector model. This means that different parts of the vector represent different contexts which are shared among sibling structures originating from the same parent text block. For example, the first part of the vector can be used to describe the document, the second part to describe the chapter, the third part to describe the paragraph and the last part to describe the individual sentence. In this thesis, we propose Hierarchical Paragraph Vectors, which exploit hierarchical document structures. When applying this novel method to sentiment analysis tasks, empirical results show that it can increase the quality of the word embeddings at the cost of greater execution overhead. i Acknowledgements I would like to thank Dr. Carsten Eickhoff for his continuous support, insightful and straightforward advice, positive criticism, proofreading, and the supervision of my master thesis. It has been a pleasure working with him. Additionally, I would like to thank Prof. Dr. Thomas Hofmann for giving me the opportunity to investigating the highly interesting topics of NLP and word embeddings, and for his valuable feedback. Furthermore, I would like to thank everyone who supported me during my studies, especially my family, my business partners and colleagues1, for their support, understanding, and positive attitude. Finally, I would like to thank Marion for her continuous positive atti-tude, proofreading, and for being by my side.

Read the paper · More papers on PaperTik