Rude-Words Detection for Indonesian Speech Using Support Vector Machine
Sashi Novitasari, Dessi Puji Lestari, Sakriani Sakti, Ayu Purwarianti · 2018
This paper presents an approach to detect rude or swear-words in Indonesian transcribed speech by using Support Vector Machine and various combinations of text and acoustic features. Rude-words considered as words which prohibited to be shown in broadcast and it usually will be censored through censorship. In the constructed framework, those words are detected by identifying the rudeness of each word of the given speech utterance. This identification aimed to be done by considering speech's context aside from the word itself, since word's rudeness related to the context of the speech. Results of the experiment show that rude-words detection which utilized textual features set that consists of word-embedding, trigram POS-tag, word list, and sentence-embedding resulted in the best performance compared to other experimented features sets. This model also outperformed the rude-words detection by acoustic model, multi-modal model, and keyword matching technique. The Fl-scores of the best model are 83.62 % in word-level detection and 87.07% in sentence-level detection.