An approach to phrase selection for offline data compression
Andrew H. Turpin, W.F. Smyth · Murdoch Research Repository (Murdoch University) · 2002
Recently several oJfline data compression schemes have been published that expend large amounts of computing resources when encoding a file, but decode the file quickly. These compressors work by identifying phrases in the input data, and storing the data as a series of pointer to these phrases. This paper explores the application of an algorithm for computing all repeating substrings within a string for phrase selection in an offiine data compressor. Using our approach, we obtain compression similar to that of the best known offiine compressors on genetic data, but poor results on general text. It seems, however, that an alternate approach based on selecting repeating substrings is feasible. Keywords: strings, of_ fline data compression, textual substitution, repeating substrings 1