Tune Your Brown Clustering, Please
Leon Derczynski, Sean Chester, Kenneth S. Bøgh · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2015
Brown clustering, an unsupervised hier-archical clustering technique based on n-gram mutual information, has proven use-ful in many NLP applications. However, most uses of Brown clustering employ the same default configuration; the appropri-ateness of this configuration has gone pre-dominantly unexplored. Accordingly, we present information for practitioners on the behaviour of Brown clustering in or-der to assist hyper-parametre tuning, in the form of a theoretical model of Brown clus-tering utility. This model is then evalu-ated empirically in two sequence labelling tasks over two text types. We explore the dynamic between the input corpus size, chosen number of classes, and quality of the resulting clusters, which has an impact for any approach using Brown clustering. In every scenario that we examine, our re-sults reveal that the values most commonly used for the clustering are sub-optimal. 1