Implementing Term Frequency-Inverse Term Frequency at Tweets in Indonesian Fraud Crime Cases
Hesty Ulfatriyani, Hanung Adi Nugroho, Indah Soesanti · 2020
Crime analysis is a methodical approach to identifying and analyzing patterns and trends in crime. Using crime analysis and text mining, we can analyze the modus operandi of crime to reduce the offenders. However, many fraud victims rarely report their problems to the police and tend to share their fraud case stories with their social media. Therefore, it needs a capable dataset to further analyze the fraud case. TFIDF as a weighting approach is used to find out the importance of a word. Thus, this study makes data derived from original data into data that can be processed for analysis. The data used are data from social media Twitter with 39,964 data that have keyword “penipuan” in Indonesian. This study uses text preprocessing techniques to clean the data from information which is not useful for the analysis process. The phases are data crawling, data cleansing, stemming, filtering, tokenizing, and visualizing data. After preprocessing data, the data will be processed into the terms frequency that appears and visualizes it. As a result, in the TF-IDF approach, the word “nomer” is the first for the word that often appear. It can be hypothesized that victims usually share their experiences of fraud that had related to the victim's personal number.