Clustering method for similar user with Miexed Data in SNS

Hyoung-Min Song, Sang-Joon Lee, Ho‐Young Kwak · Journal of the Korea Society of Computer and Information · 2015

The enormous increase of data with the development of the information technology make internet users to be hard to find suitable information tailored to their needs. In the face of changing environment, the information filtering method, which provide sorted-out information to users, is becoming important. The data on the internet exists as various type. However, similarity calculation algorithm frequently used in existing collaborative filtering method is tend to be suitable to the numeric data. In addition, in the case of the categorical data, it shows the extreme similarity like Boolean Algebra. In this paper, We get the similarity in SNS user's information which consist of the mixed data using the Gower's similarity coefficient. And we suggest a method that is softer than radical expression such as 0 or 1 in categorical data. The clustering method using this algorithm can be utilized in SNS or various recommendation system.            2013년 한 해 전 세계에서 생산된 데이터 총량은 4.4ZB에 달하고 2020년에는 10배 증가하여 44ZB에 이를 것으로 예측되고 있다[1]. 이런 엄청난 양의 데이터로 인하여 정보선택의 문제가 발생하고 있으며, 사용자에게 불필요한 내용을 제거하고 필요한 정보만을 제공해주는 정보 필터링 기법과 여러 곳에 흩어진 콘텐츠를 하나의 콘셉트나 주제로 모아서 보여주는 콘텐츠 큐레이션 기법이 유행하고 있다[2][3].일반적으로 추천시스템에 사용되는 정보 필터링기법으로는 협업적 필터링과 내용기반 필터링이 있다. 특히 협업적 필터링에서는 사용자와 성향이 유사한 사용자를 찾기 위해 유사도를 계산해야 하는데 기존 시스템에서 많이 사용되어왔던 피어슨 상관계수(Pearson correlation coefficient), 코사인 유사도(Cosine Similarity) 등의 유사도 계산법은 수치형 데이터에만 활용이 가능하다. 그러나 SNS(Social Network Service) 상에 존재하는 사용자 정보는 수치형 데이터뿐 아니라 다양한 형태의 데이터로 이루어져있기 때문에 혼합형 데이터의 유사도를 구하는 방법이 필요하다. 본 논문에서는 Gower 유사도 계수(Gower's similarity coefficient)를 사용하여 수치형과 범주형의 혼합형 데이터로 이루어진 데이터의 유사도를 구하고 범주형 데이터의 경우 부울대수 형태의 극단적인 표현이 아닌 완화된 형식으로 계산하는 방법을 제안한다. 본 논문에서는 2장에서 관련 연구에 대하여 설명하고, 3장에서는 제안하는 알고리즘에 대하여 설명하였으며, 4장에서는 결과에 대한 분석을 실시하고, 5장에서 결론을 기술하였다.

Read the paper · More papers on PaperTik