DiscoveringHigh Order Models in Evolving Data
Shuigeng Zhout, Philip S. Yu · 2008
Manyapplications aredriven byevolving data - patterns inwebtraffic, programexecution traces, network event logs, etc., areoften non-stationary. Building prediction models forevolving databecomesanimportant andchallenging task. Currently, mostapproaches workbychasing trends, thatis, theykeeplearning orupdating modelsfromtheevolving data, andusethese impromptu modelsforonline prediction. Inmany cases, this proves tobebothcostly andineffective - muchtime iswasted onre-learning recurring concepts, yettheclassifier may remainonestepbehindthecurrent trendallthetime. Inthis paper, wepropose tominehigh-order modelsinevolving data. Moreoften thannot,therearealimited numberofconcepts, orstable distributions, inthedatastream, andconcepts switch between eachother constantly. We mineallsuchconcepts offline froma historical stream, andbuildhighquality modelsfor eachofthem.Atruntime, combining historical concept change patterns andcuesprovided byanonline training stream, wefind themostlikely current concept anduseitscorresponding models toclassify datainanunlabeled stream. Theprimary advantage of thehigh-order modelapproach isitshighaccuracy. Experiments showthatinbenchmark datasets, classification error ofthehigh- ordermodelisonlyasmallfraction ofthatofthecurrent best approaches. Another important benefit isthat, unlike state-of-the- artapproaches, ourapproach doesnotrequire users totuneany parameters toachieve asatisfying result onstreams ofdifferent characteristics.