Thai Tokenizer Invariant Classification Based on Bi-LSTM and DistilBERT Encoders
Noppadol Kongsumran, Suphakant Phimoltares, Sasipa Panthuwadeethorn · 2022 19th International Conference on Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology (ECTI-CON) · 2022
This research aims to solve Thai word boundaries problem by grouping vectors of words obtained from various tokenizers in the same sentence using Bi-LSTM and DistilBERT encoders. Generally, one sentence can achieve different sets of words from multiple tokenizers, resulting in uncontrollable accuracies in various tasks. The proposed method based on Bi-LSTM and DistilBERT is to train and transform data with triplet hard loss to a particular domain where the embedding vectors of each same sentence are much closer to each other. In this study, sentiment classification is conducted to show that our method can be used as pre-trained process for other tasks.