Automatic rhetorical sentence categorization on Indonesian meeting minutes
Ghoziyah Haitan Rachman, Masayu Leylia Khodra · 2016
Meeting minutes contains much important information of meeting. Since meeting minutes is unstructured document, in order to easily get and summarize this information, classification for every sentence in meeting minutes should be conducted. Some works in this research area has been done for meeting minutes in English, but not conducted yet in Indonesian. Therefore, this paper aims to present the rhetorical sentence categorization from Indonesian meeting minutes by utilizing some features, i.e. length, position, previous label, significant terms, and cue phrases per class. Then, this paper shows the result of employing SMOTE and resampling for balancing the existing instances per class. Every experiment is tested in four classifiers, namely Naïve Bayes, SVM Linear, IBk, and J48 tree. It shows that the use of previous label and both of significant term and cue phrase per class improves performance. Then it shows also that resampling is better than SMOTE. After doing the 10-fold cross-validation in IBk classifier, model using SMOTE achieved F-measure of 85.22% and resampling model achieved 94.52%.