Hash-Grams On Many-Cores and Skewed Distributions
Edward Raff, Mark McLean · 2018
When using n-grams for features, it is often the case that an expedient and effective first-pass of feature selection can be performed by picking the top-k most frequent features. The hash-gram approach was introduced as a method of quickly performing this feature selection. In this work we identify a failure case of parallelizing the hash-gram algorithm to a large number of CPU cores P when the data is highly skewed. We resolve this issue to produce a hash-gram algorithm with consistent performance across potential skewness-es and number of CPU cores, making it practically usable for big-data cases where more powerful compute is needed.