Automatic keyword extraction: An ensemble method
Tayfun Pay, Stephen Lucci · 2017
We design and analyze an ensemble method for automatically extracting keywords from single documents. The automatic keyword extractors that we use in our approach are: TextRank [1] RAKE [2] and TAKE [3]. Each one of these automatic keyword extractors provides a set of candidate keywords for the ensemble method. Our approach then prunes this set of candidate keywords by applying a filtering heuristic and then recalculates their scores according to prescribed metrics. Then, dynamic threshold functions are applied to select a set of keywords for a given document. We used the data set in [4] to test the accuracy of our approach. We obtained a better overall performance when compared to each one of the individual keyword extractors that we used in constructing our ensemble method.