Preprocessing Role in Analyzing Tweets Towards Requirement Engineering
Amir Mohammad Ebrahimi, Ahmad Abdollahzadeh Barforoush · 2019
Twitter is an important source for crowdsourcing software requirements and can be used as a communication channel between technical stakeholders and end users of software products. Previous studies have shown that software users share their opinions about the software through short messages called tweets. These tweets might contain worthwhile information (such as bug reports and user requests) which relates to the software requirements. Machine learning and Natural language processing approaches have been applied to elicit these technical tweets. Particularly, when learning techniques are considered, the success rate of the preprocessing process to overcome shortcomings in analyzing tweets becomes crucial. Therefore, in this paper, we propose a method which involves more than 6000 experiments in order to systematically analyze the effects of eleven preprocessing techniques on the classification performance of requirements-related tweets. We train and test three popular supervised learning algorithms on a dataset of 4000 tweets which is a sample of twitter feeds. Finally, we analyze our evaluation results from three different viewpoints to answer the research question. Our findings have demonstrated the crucial role of preprocessing, and also are feasible to provide more insights in designing approaches which can deal more accurately in case of elicitation such tweets.