A Study of Criterion for Efficient Clustering Estimation of Temporal Data

Jin-Ho Jeon, Minsoo Kim · The Journal of the Institute of Webcasting, Internet and Telecommunication · 2011

Abstract Most real world system such as world economy, management, medical and engineering applications contain a series of complex phenomena. One of common methods to understand these system is to build a model and analyze the behavior of the system. As a first step, Determining the best clusters on data. As a second step, Determining the model of the cluster. In this paper, we investigated heuristic search methods for efficient clustering. It is also confirmed that the Bayesian Information Criterion more reliable than Cheeseman-Stutz ones. Key Words : Temporal 데이터, 군집, 기준, 한계우도 * 정회원, 관동대학교, 경영학과 ** 정회원, 관동대학교, 호텔경영학과 접수일자 2011.6.30, 수정일자 2011.9.19게재확정일자 2011.10.14 Ⅰ. 서 론 실세계의 많은 정보시스템들은 동적인 특징을 가진다. 즉, 시간적인 특징들에 의해서 묘사되고, 그것들의 값들은 관측기간 동안 의미 있게 변함을 의미한다. 이렇게 시간의 흐름에 따라 발생한 데이터를 수집하여 기록한 것을 temporal 데이터라 한다. temporal 데이터 내에 내재하는 속성들, 물품 또는 사건들을 통해 연관성 또는 순차 패턴과 같은 특징이 명확한 분야의 연구는 많이 진행되어왔다 [1] . 이러한 temporal 데이터를 분석하여 내포하고 있는 특징을 찾아낸다면 그러한 특징들을 통하여 temporal 데이터를 이해하고 분석하는데 도움을 줄 것이다.대용량의 temporal 데이터의 분석을 위한 과정은 두 단계로 살펴볼 수 있다. 첫 번째 과정은 발생되어진 대용량의 temporal 데이터들을 유사한 데이터 객체들끼리의 군집화 과정이다. 두 번째 과정은 각 군집을 잘 설명할 수 있는 적합한 모델을 생성, 학습하는 과정이다.본 연구에서는, 위의 두 단계 과정 중에서 첫 번째 과정, 즉, temporal 데이터의 군집화 과정에서 데이터들의 특징을 표현하기에 적합한 군집 수를 추정하는 기준에

Read the paper · More papers on PaperTik