Verb Phrase Extraction in a Historical Context
Eva Pettersson, Beáta Megyesi, Joakim Nivre · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2014
In recent years, large volumes of historical text have been made digitally available. There is however still a lack of suitable language technology tools for exploring these texts in an automated way. In this paper we present a method for automatic identification and extraction of verbs and their complements in historical text. This work has been carried out in cooperation with historians in the context of the Gender and Work project (GaW), where researchers are building a database with information on what men and women did for a living in the Early Modern Swedish society (approx. 1550–1800) (Agren et al., 2011). Currently, historians are manually going through historical documents, searching for relevant text passages to store in the database. In this process, it has been noticed that working activities often are described in the form of verb phrases, such as chop wood, sell fish or serve as a maid. An interesting language technology challenge is thus to try to automatically extract verb phrases from historical text, and present these to the historians as a list of candidate phrases for database inclusion. Ultimately, such a tool would enable the historians to fill the GaW database with relevant phrases in a shorter period of time. In Section 2 we present our proposed method, whereas the data used in our experiments are introduced in Section 3. The approaches we use for spelling normalisation are described in Section 4. Finally, results are given in Section 5, while conclusions are drawn in Section 6.