Efficient Clustering Technique for University Admission Data
Abdul Fattah, Maher Zuhair Mohammed, Philip S. Yu, Tarek F. Gharib · 2012
ABSTRACT Educational Data Mining (EDM) is the process of converting raw data from educational systems to useful information that can be used by educational software developers, students, teachers, parents, and other educational researchers. In this paper, we present an efficient clustering technique for King Abdulaziz University (KAU) admission data. The model uses K-Means algorithm. The clustering quality is evaluated using the DB internal measure. Experimental results show that K-Means achieves the minimum DB value that gives the best fits natural partitions. Additional analysis is also presented from the perspective of university admission office. Keywords Educational Data Mining (EDM); Data Clustering; University Admission Data; Clustering Evaluation. 1. INTRODUCTION Data mining aims at the discovery of useful information from large collection of data. Recently, there are increasing research interests in using data mining in education. This new emerging field, called Educational Data Mining (EDM), concerns with developing methods that discover knowledge from data originating from educational environments [1]. Educational data mining (EDM) differs from knowledge discovery in other domains in several ways. One of them is the fact that it is difficult, or even impossible, to compare different methods or measures a posteriori and decide which is the best. Take the example of building a system to transform hand-written documents into printed documents. This system has to discover the printed letters behind the hand-written ones. It is possible to try several sets of measures or parameters and experiment what works best. Such an experimentation phase is difficult in the education field because the data is very dynamic, can vary a lot between samples and teachers just cannot afford the time and access to the expertise to do these tests on each sample, especially in real time. Therefore, one should care about the intuition of the measures, parameters or methods used in educational data mining [2]. Clustering algorithms attempt to organize unlabeled input vectors into clusters or “natural groups” such that points within a cluster are more similar to each other than vectors belonging to different clusters [3]. Clustering has been used in exploratory pattern-analysis, grouping, decision-making, and machine-learning situations, which include data mining, document retrieval, image segmentation, and pattern classification. The clustering methods are of five types: hierarchical clustering, partitioning clustering, density-based clustering, grid-based clustering, and model based clustering [4]. Each type has its advantages and disadvantages. Verma et al. Provided a comparative study of commonly used clustering algorithms in data mining field [5]. Their study compared the performance of six types of clustering techniques: K-Means, Hierarchical, DBScan, Density-based, OPTICS and EM algorithms. Their experimental results showed that K-Means were faster and achieved good clustering results than other algorithms but it was sensitive to noise (if exists). In this paper, we present an efficient clustering model for King Abdulaziz University (KAU) admission data. The model uses K-Means algorithm and DB measure as internal clustering quality evaluation index. The rest of this paper is organized into five sections. In section 2, the clustering model and algorithms are briefly reviewed. Section 3 presents the KAU admission system as a case study. In section 4, experimental results are presented and analyzed with respect to model results and admission system perspective. Finally, the conclusions of this work are presented in Section 5.