Efficient Data Structures for MassiveN-Gram Datasets
Giulio Ermanno Pibiri, Rossano Venturini · 2017
The efficient indexing of large and sparse N-gram datasets is crucial in several applications in Information Retrieval, Natural Language Processing and Machine Learning. Because of the stringent efficiency requirements, dealing with billions of N-grams poses the challenge of introducing a compressed representation that preserves the query processing speed.