From Text to Data: On The Role and Effect of Text Pre-Processing in Text Mining Research
Matthias Rüdiger, David Antons, Torsten Oliver Salge · Academy of Management Proceedings · 2017
Over the last decade, text mining has established itself as a valuable method for structuring the content of large text corpora such as journal and newspaper articles or patent applications. Despite the broad acceptance of text mining, however, some of its methodological challenges are yet to be explored. In particular, the role and effects of text pre-processing necessary to convert text to data suitable for subsequent mining have remained largely unexamined. Even more, text mining studies rarely document the precise text pre- processing steps performance or the sensitivity of their results to changes in text pre-processing. This not only limits the validity and reliability of their studies, but also renders subsequent replication attempts difficult at best. It is against this backdrop that we systematically explore the role and effect of text pre-processing in the context of text mining research. For this purpose, we review the methodological state-of-the-art and develop an experimental testing procedure to quantify the effect of text pre-processing within a realistic use case. Our work helps researchers assess the sensitivity of their results to changes in the text pre-processing procedure and equips them with best practice guidelines on how to conduct and report text pre-processing.