From Linguistic Descriptions to Language Profiles.

Shafqat Mumtaz Virk, Harald Hammarström, Lars Borin, Markus Forsberg, Søren Wichmann · Data Archiving and Networked Services (DANS) · 2020

7th Workshop on Linked Data in Linguistics (LDL-2020).Building tools and infrastructuresPast years have seen a growing interest in the application of knowledge graphs and Semantic Web technologies to language resources, and their publication as linked data on the Web.As of today, a large amount of language resources were either converted or created natively as linked data on the basis of data models specifically designed for the representation of linguistic content.Examples are wordnets, dictionaries, corpora, culminating in the emergence of a Linguistic Linked Open Data (LLOD) cloud (http://linguistic-lod.org/).Since its establishment in 2012, the Linked Data in Linguistics (LDL) workshop series has become the major forum for presenting, discussing and disseminating technologies, vocabularies, resources and experiences regarding the application of semantic technologies and the Linked Open Data (LOD) paradigm to language resources in order to facilitate their visibility, accessibility, interoperability, reusability, enrichment, combined evaluation and integration.The LDL workshops contribute to the discussion, dissemination and establishment of community standards that drive this development, most notably the OntoLex-lemon model for lexical resources, as well as standards for other types of language resources still under development.The workshop series is organized by Open Linguistics, founded 2010 as a Working Group of the Open Knowledge Foundation 1 with close involvement of related communities, such as W3C Community Groups, and international research projects.It takes a general focus on LOD-based resources, vocabularies, infrastructures and technologies as means for managing, improving and using language resources on the Web.As technology and resources increasingly converge towards a LODbased ecosystem, this year we particularly encouraged submissions on Linked-Data Aware Tools and Services and Linked Language Resources Infrastructure, i.e. managing, curating and applying LLOD technologies and resources in a reliable and reproducible way for the needs of linguistics, NLP and digital humanities.After ten years of community work, a critical mass of LLOD resources is already in place, yet, there is still a need to develop a robust ecosystem of tools that consume linguistic linked data.Recently started research networks and European projects are working in the direction of building sustainable infrastructures around LRs, with linked data as one of the core technologies.LDL-2020 is thus supported by the COST Action "European network for Web-centred linguistic data science" (NexusLinguarum) and two Horizon 2020 projects, the European Lexicographic Infrastructure (ELEXIS), and Prêt-à-LLOD, which focuses on providing an infrastructure for linguistic data to be ready to use by state-of-the-art technologies.With a focus on building tools and applications, the 7th Workshop on Linked Data in Linguistics (LDL-2020) was organized in conjunction with the 12th Language Resource and Evaluation Conference (LREC-2020).We received a total of 23 submissions out of which 12 were accepted (acceptance rate 52%).Due to Covid-19, LDL-2020 was not taking place as a physical meeting, but as a virtual event 2 .Presentations of the accepted papers were organized in three groups with four presentations each, on modelling, applications and lexicography, respectively. Lexicography Abgaz describes on-going work onUsing OntoLex-Lemon for Representing and Interlinking Lexicographic Collections of Bavarian Dialects, comprising two main components, a questionnaire with details about questions, collectors, paper slips etc., and a lexical dataset which contains lexical entries (answers) collected in response to the questions.The paper describes how the original TEI/XML format is transformed into Linguistic Linked Open Data to produce a lexicon for Bavarian Dialects.With Linguistic Linked (Open) Data and, especially, the OntoLex vocabulary now being widely adapted throughout lexicography, there is a demand for tools, both for exploiting linked lexical data and for creating a user-friendly access to it.In Involving Lexicographers in the LLOD Cloud with LexO, an Easy-to-use Editor of Lemon Lexical Resources, Bellandi and Giovannetti describe LexO, a collaborative web editor of OntoLex-Lemon resources.As for tools for lexicography, Gun Woo Lee et al. describe Supervised Hypernymy Detection in Spanish through Order Embeddings, based on a hypernymy dataset for Spanish built from WordNet and the use of pretrained word vectors as input.Finally, Nielsen reports on Lexemes in Wikidata, i.e., the way that Wikidata records data about lexemes, senses and lexical forms and exposes them as Linguistic Linked Open Data and the growth and development of this data set since its first establishment in 2018.

Read the paper · More papers on PaperTik