XNLRDF, an Open Source Natural Language Resource Description Framework

Oliver Streiter, Mathias Stuflesser · Institutional Repositories DataBase (IRDB) · 2005

XNLRDF represents an unseen attempt to collect, formalize and formally describe language resources on a large scale so that they can be used automatically by computer applications.XNLRDF is intended to become a free software distributed in XML-RDF.This software is designed to be accessed by computer applications like Web-browsers, mail-tools, Web-crawlers, information retrieval (IR) systems or Computer Assisted Language Learning (CALL) systems.It proposes to replace idiosyncratic ad-hoc solutions for Natural Language Processing (NLP) tasks by a standard interface to XNLRDF.The linguistic information in XNLRDF covers a wide range of written languages and extends the information offered by Unicode so that basic NLP tasks like language recognition, tokenization, stemming, tagging, term-extraction etc can be performed.With more than 1.000 languages used in the Internet and their number continually rising, the design and development of such a software becomes a pressing need.In this paper we introduce the basic design of XNLRDF, the type of information the first prototypes will provide and describe the current state of the project. 1.XNLRDF as a Natural Extension of Unicode 1.1.Advantages of UnicodeWith the advancement of Unicode, the processing of many languages, for which previously specific techniques were required, has become simplified.Unicode describes characters of language scripts by giving them a unique code point and properties like uppercase, lowercase, decimal digit, mark, punctuation, hyphen, separator or the script.Operations on the characters such as uppercasing, lowercasing and sorting are defined as well.Any computer application which is not endowed with particular linguistic knowledge is thus better off when processing texts in Unicode than in traditional encodings such as latin1, big5 or koi-r.With Unicode, the recognition of words, numbers and sentences may be performed without additional resources for many languages.Accordingly, various libraries for programming languages like Java, C++ or C have been developed to grant the programmer easy access to the information contained in Unicode.

Read the paper · More papers on PaperTik