Automatic Domain-Specific Corpora Generation from Wikipedia - A Replication Study
Seniru Ruwanpura, Cale Morash, Momin Ali Khan, Adnan Ahmad, Gouri Ginde · 2023
Replication studies help mature our knowledge and attempt to validate the findings of a prior piece of research. However, these studies are still rare in the Requirements Engineering field. Additionally, the rapidly advancing realm of Natural Language Processing (NLP) is creating new opportunities for efficient, machine-assisted workflows application which can bring new perspectives and results to the forefront. Thus, in this paper, we replicate and extend a previous study (baseline), a tool, WikiDoMiner, which automatically generated domain-specific corpora by crawling Wikipedia. In this study, we investigated and executed the implementation of WikiDoMiner (open-sourced code from the original paper) to recreate the results. This allowed us to strengthen the external validity of the original study. We extended the baseline to evaluate additional data sets and generated nuanced results using state-of-the-art NLP techniques such as Bidirectional Encoder Representations from Transformers (BERT). Results showed that due to the growing content in Wikipedia, the corpus generated for the Railways and Networks domains did not precisely match the results from the baseline. However, utilizing the state-of-the-art KeyBERT library from the Huggingface AI community enhanced the results, eventually generating a meaningful corpus compared to the baseline.