LINGUISTICS IN THE INTERNET AGE: TOOLS AND FAIR USE

William Dodge Lewis, Scott O. Farrar, D. Terence Langendoen · 2006

The current work explores the fair use of linguistic data in the context of the Internet. It is argued that the field is now at a critical turning point with respect to the way data are reused and disseminated, since current copyright laws are not adequate to protect digital language data. The problem of data misuse does not only originate from the fact that data are now broadly accessible to anyone with an Internet connection, but is exacerbated by ever more precise search engines that may access and reuse data in an automated fashion. The issue of what constitutes fair use in this domain, especially when considering the boundaries established under copyright law, is neither clear nor well-defined. As a solution, a set of principles, called Principles of Reuse and Enrichment of Linguistic Data, or PRELDs, is proposed for data found on the Internet. Each of these principles is developed by considering examples from current research that show how linguists have used, reused, and, at times, misused data in traditional print media. The principles are put to use by considering how automated linguistic tools, especially those that have the potential to blur the lines of fair use, can be made to promote linguistics as a cutting-edge scientific enterprise, while preserving the rights of authors and respecting individual scholarship.

Read the paper · More papers on PaperTik