On-Line Unsupervised Learning for Information Compression and Similarity Analysis of Large Data Sets

Gancho L. Vachkov, Hidenori Ishihara · 2007

The growing huge amount of information from the operations of complex processes and systems requires suitable methods for information compression. Therefore in this paper three unsupervised learning algorithms for information compression are proposed and analysed, namely the fixed-model learning (FML), the growing-model learning (GML) and the on-line model learning (OML) algorithms. They convert the original large data set into a much smaller set of neurons in the same dimensional space. It is shown that the OML algorithm is the fastest one and the most suitable for large data compression. A procedure for similarity analysis of the compressed models is also presented and illustrated in the paper. It uses the preselected Key Points from the compressed model for comparison.

Read the paper · More papers on PaperTik