Hidden Structures: Clustering, String Distance, Text Vectors and Topic Modeling

Ted Kwartler · 2017

This chapter first gives tools to identify underlying structure in text. These include document clustering, string distance calculations and topic modeling techniques. The chapter also gives numerous approaches to clustering documents. Each has different approaches and popularity. Methods of clustering documents include k-means clustering, and k-mediod clustering. K-mediod and k-mean approaches are often similar but it is worthwhile to explore many approaches during the course of an analysis. The chapter follows the six-step text mining process for document clustering in an HR analytics case study. It further discusses work to find the distances between strings in a different manner than Euclidean distance or cosine similarity. When performing string distance analysis it is important to know the method with which the strings were collected, the ways distance calculations impact results and subsequently the resulting cluster analysis. Finally, the chapter examines a relatively new R package called text2vec.

Read the paper · More papers on PaperTik