D16.4: Final Report on Natural Language Processing
Andreas Vlachidis, Douglas S. Tudhope, M. Wansleeben, Janet Azzopardi, Green, Katie, Lei, Xia, Wright, Holly · UCL Discovery (University College London) · 2017
This document is a deliverable (D16.4) of the ARIADNE project (“Advanced Research Infrastructure for Archaeological Dataset Networking in Europe”), which is funded under the European Community's Seventh Framework Programme. It presents the final results of the work carried out in Tasks 16.2 “Natural Language Processing (NLP)”. NLP is an interdisciplinary field of computer science, linguistics and artificial intelligence that uses many different techniques to explore the interaction between human (natural) and computer languages. The partners continued to focus on one of the most important, but traditionally difficult to access resources in archaeology; the largely unpublished reports generated by commercial or “rescue” archaeology, commonly known as “grey literature”, exploring both rule-based and machine learning NLP methods, the use of archaeological thesauri in NLP, and various Information Extraction (IE) methods in their own language. USW extended their English language rule based methods using the GATE toolkit for NER (Named Entity Recognition) to Dutch and Swedish language grey literature reports, in collaboration with LU and DANS (Dutch reports) and SND (Swedish reports). This made use of glossaries and thesauri from the partners, including the Dutch Rijksdienst Cultureel Erfgoed (RCE) Thesauri. The process of importing the thesauri resources into a specific framework (GATE), and the suitability and performance of the selected resources when used for the purposes of Named Entity Recognition (NER) were analysed. The NER techniques were focused on the general archaeological entities of Archaeological Context, Material, Physical Object (Finds), Monument, Place, and Temporal (Time Appellation). The methods proved capable of extracting CIDOC CRM element and in some case studies Getty Art and Architecture Thesaurus concepts, in addition to the native vocabularies. General archaeological NLP (GATE) pipelines for English, Dutch and Swedish have been developed. In addition experimental pipelines were developed for two exploratory thematic case studies on data integration, where the output is expressed as RDF Linked Data via a CRM based data model. An English language pipeline is available for a numismatic case study. English, Dutch and Swedish pipeline are available for a case study of item level data/NLP integration on a loose theme based around archaeological interest in wooden objects and their dating, as expressed in different kinds of datasets and reports. Both case studies have resulted in interactive demonstrators operating over the ARIADNE Linked Data Cloud. All 7 pipelines are freely available as open source ARIADNE outcomes. The Archaeology Data Service (ADS) at the University of York continued developing a machine learning-based NLP technique which has now been integrated it into a new metadata extraction web API, which takes previously unseen English language text as input, and identifies and classifies named entities within the text. The outputs can then be used to enrich resource discovery metadata for existing and future resources. This API can be incorporated into existing interfaces and used by archaeological practitioners to automatically generate metadata related to text-based content uploaded on a per-file basis, or by using batch creation of metadata for multiple files. This report presents the final results of the work carried out to date, and presents the issues to be addressed during the remainder of the ARIADNE Project