Research on Classification of Kazakh Questions Integrate with Multi-feature Embedding

Gulizada Haisa, Gulila Altenbek, Hayinaer Aierzhati, Kaden Kenzhekhan · 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT) · 2021

Kazakh is an agglutinative language, and this feature results in data sparseness to some extent. In addition, the Kazakh question sentences lack strict grammatical rules, and the word forms are changeable and irregular. Due to this, this article proposes a CNN+BiGRU question classification model and attention mechanism that integrate multi-features based on the Kazakh language characteristics. Kazakh words and language features are used as input for the neural network. The CNN network generates high-dimensional semantic features and transmits the output to the BiGRU network. The BiGRU layer models the context information, then the Attention layer concentrates on the input features, filters out the unnecessary information, and completes the classification with SoftMax. The research in this paper shows that our model effectively integrates language features, avoids data sparsity, improves the model's performance during training, and has higher performance on the classification of Kazakh questions.

Read the paper · More papers on PaperTik