A NEW SIMILARITYMEASURE FOR MICROARRAY DATA ANALYSIS
Hong Yan · 2005
A numberofclustering algorithms havebeenusedfor micoarray dataanalysis. However, theperformance of thesemethodsissignificantly degraded duetothe presence ofnoise. Inthis paper, we introduce arobust clustering algorithm basedonanewsimilarity measure. Thekeyconcept ofthenewsimilarity measure isto measure thesimilarity between twodatapoints bytheir sub-dimensions. Forexample, assumethatxi, x2andX3 are10dimensional data vectors. Thedata point X3issaid tobecloser toxlthanx2ifmorethanhalfofthe dimensions ofxlandX3arecloser toxlthanx2.Thus, if twopatterns areverysimilar except asmall amountof features ornoise, this measure will preserve thesimilarity. Experimental results showthat theclustering algorithm using this measure produces better results thancommonly usedsimilarity measures.