NAFIPS 2005-2005Annual Meeting oftheNorth American Fuzzy Information Processing Society

A ModifiedCompetitive Agglomeration · 2005

Clustering algorithms areinvaluable methods for organizing dataintouseful information. TheCARDAlgorithm (11isonesuchalgorithm thatisdesigned toorganize user sessions intoprofiles, whereeachprofile wouldhighlight a particular typeofuser.TheCARD algorithm isa viable candidate forwebclustering. However itdoeshavelinitations suchaslongexecution time. Inaddition, thedatapreparation for thealgorithm's requirements employs concepts thatare incomplete. Theselimitations ofthealgorithm will beexplored andmodified toyield amorepractical andefficient algorithm. I.INTRODUCTION AstheWorldWideWebhasgrowninsize, animmense amountofdatahasbeenproduced.Ifthisdataisnot retrieved andorganized correctly thentheinformation that the datacanprovide isessentially wasted. Usersoften become frustrated because theyarecertain thattheinformation that theyarelooking forisonaparticular website, butthewebsite issolarge that itisalmost impossible tofindtheinformation. Inthiscase, anadaptive website canbecomeveryhelpful. Thiswebsite wouldfollow theclickstream oftheuser.That is, thewebsite will keepaccount ofwebpages that theuseris looking at,andrecommend other pagesthat previous users found useful. Thewebsite will implement this adaptability byhaving a database oftypical userprofiles. A userprofile isdefined as anabstract modelthatsummarizes therelevance ofeach URLonasite relative toagroupofuserssharing asimilar interest (2).Theuserprofile willbecharacterized by previous usersessions. A usersession isasetofwebpages thata userexamined within a specified timeperiod. To organize theusersessions intoprofiles, thewebsite administrator could examine eachusersession andmanually place theminto profiles. However, this isunrealistic formany reasons, forinstance, manywebsites havemillions ofusers accessing itdaily anditisimpossible tofind patterns bymere manualexamination. Therefore, wecaneasily seethat an unsupervised clustering algorithm isanattractive choice for clustering usersessions into profiles ofseveral types oftypical users since itrelies onuseraccess patterns andiscapable of examining large amounts ofdata inafairly reasonable amount oftime(1).

Read the paper · More papers on PaperTik