Duplicate Question Management and Answer Verification System
Somak Mukherjee, N S Kumar · 2019
Management of large data sets of question papers can be cumbersome, especially when dealing with potential duplicate or erroneous questions. The addition of a natural language system that automatically handles these issues would greatly speed up the verification of such data sets. This is a tool for identifying semantic similarity between sentences in plain-text English. Handpicked features were selected which included simple structural features and word embedding features using word2vec with multiple distance metrics between the resulting sentence vectors. The model is trained on weak hardware allowing for sufficiently high accuracy on low end machines. Results demonstrate the effectiveness of boosting for improving the performance of simple learning models, allowing for complex learning in the absence of high end hardware.