K-Means algorithm and modification using gain ratio
Ryan Dhika Priyatna, Tulus Tulus, Marwan Ramli · IOP Conference Series Materials Science and Engineering · 2018
K-Means is one method in data mining that can be used to perform grouping clustering of data.Accurate data processing can be done by processing the data source.Each collection or data warehouse can provide important knowledge into valuable information, constraints on this method, if the cluster point is chosen randomly so that the resulting data may vary, if the value is not good, then the resulting grouping is less than optimal.Furthermore, failure to outliers in the process of grouping data include determining whether a data item is an outliers of a cluster of course and whether small amounts of data form a separate cluster.where the gain ratio is used to calculate the attribute's influence on the target of a data gain ratio is the development of the information gain, where the gain ratio eliminates the bias value of each attribute.The result of the research is to calculate the weights in each attribute by using the gain ratio and make the modeling and classification into the method of K-means.