Research on Financial Fraud Text Classification Based on PET-BiLSTM
Feifei Zou, Su Juan Hu, Yu Hang Wei, Zejun Yan, Sijun Chan · 2024
To address challenges associated with obscure features, unbalanced data, and diverse fraud types in Internet financial texts, this paper proposes the PET-BiLSTM financial fraud text classification model to enhance the semantic understanding and reasoning of financial fraud texts. Crawling news, comments and short messages about financial fraud from open source websites, desensitizing and cleaning the texts, labeling them automatically through LDA topic distribution model and financial fraud thesaurus, constructing prompt templates containing mask positions, forming cloze sentences for each sample, and transforming them into multi-classification tasks. The pre-trained BERT model is used to learn semantic information, combined with bidirectional LSTM (BiLSTM) to extract fraudulent text features, and supervised training is carried out. Experimental results show PET-BiLSTM achieving 89.11% and 84.54% F1 values in multi-classifying fraud types and three-classifying fraud degree identification tasks, surpassing deep learning baselines by 4.21% to 15.35% and 2.21% to 9.12%, respectively. PET-BiLSTM also outperforms ChatGPT in zero-shot prompt learning by 67.11% and 48.54%. The model demonstrates superior performance in understanding knowledge facts and fraud reasoning in the financial fraud text classification task, effectively addressing complex text fraud detection challenges within the financial field.