On Natural Language Processing Applications for Military Dialect Classification
Charith Gunasekara, Tobias Carryer, Matt Triff · 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) · 2021
The authors explore potential Natural Language Processing (NLP) models for the classification of military text. A dataset of military articles were compiled by web scraping public access military documents published by military branches in Australia, Canada, the UK and USA, using web crawler algorithms. Various NLP algorithms were tested and compared to evaluate the classification performance on the military documents to identify the strengths and weaknesses of each model. The top performing model, multiclass BERT, achieved 99.9% accuracy for classification when predicting the military branch for a withheld test set of articles. However BERT did not perform as accurately in the binary classification models when trained on military documents. Overall model performances, training time and the effect of input text length on model performance is compared across all models.