Identifying Data Set Texture using Normalized Compression Distance

Shahany Habeeb, Syam Gopi · International Journal of Engineering Trends and Technology · 2015

Models that do not preserve text structure or that preserve text structure can be used for presenting text data sets. Here the main hypothesis is that depending on the nature of data set, there can be advantages of using a model that preserves text structure over one that does not, and vice versa. The key is to determine the best way of presenting a particular data set, based on the data set itself. Here different distortion techniques are analyzed on the bases of compression distances for identifying texture of data sets.

Read the paper · More papers on PaperTik