Efficient Resume Classification Through Rapid Dataset Creation Using ChatGPT
Panagiotis Skondras, George Psaroudakis, Panagiotis Zervas, Giannis E. Tzimas · 2023
In this article, is presented a study on the problem of resume classification using machine learning algorithms. The use of machine learning algorithms in natural language processing (NLP) has become increasingly popular as they can accurately classify resumes, saving organizations time and money in the hiring process. The effectiveness of these algorithms relies heavily on the quality and quantity of data used for training neural network models. The choice of the underlying model architecture is crucial for achieving good performance as well as a well-curated dataset is needed for effective training. This is especially true for classification tasks, where accurate annotations are required to achieve high performance. We used the Open AI API [1], in order to generate structured and unstructured resumes based on specific criteria. The generated resumes were cleaned, annotated and used to train a feedforward neural network which relied on Universal Sentence Encoder (USE) [2] embeddings. This approach allowed us to efficiently generate large amounts of annotated data that are tailored to specific occupation categories for our classification task. To evaluate the effectiveness of our model, a dataset of resumes was constructed from indeed.com. This approach ensures that our model is not biased towards working only on the generated resumes but it can also perform very well on real-world data. The model identified the relevant occupation category per resume and achieved an overall accuracy of 98% on the ChatGPT -generated test set and 85% on the indeed.com dataset.