Identifying Well-formed Questions using Deep Learning
Navnoor Chhina · UVic’s Research and Learning Repository (University of Victoria) · 2020
Deep Learning is dominant in the field of Natural language processing, thanks to its performance over the statistical methods. The high performance is driven by the recent transfer learning approaches, where we pre-train a language model over a large corpus and use the pre-trained model for fine-tuning on any specific task. Recent advancements in transfer learning suggest how using the pre-trained model and fine-tuning for a few training steps can give us state-of-the-art results. Therefore, in this project for the classification of well-formed natural language questions, we use both the learning of a model from scratch and transfer learning. We can use the pre- trained base models from BERT, ALBERT and XLNet and add a simple classifier layer, which gives us better results than learning the model from scratch in very few epochs. We also sample a subset of classified queries by our model and run the queries on the Google search engine and confirm that using a model that can identify the well-formedness of the queries would be helpful to the search engine by reducing the downstream compounding errors for the natural language processing pipeline.