Finnish dialect atlas for quantitative studies
Sheila M. Embleton, Eric S. Wheeler · Journal of Quantitative Linguistics · 1997
Before we can do quantitative studies of large volumes of dialect material (such as Embleton & Wheeler, 1996, on English), it is necessary to have machine‐readable sources of data. For Finnish dialects, the principal source of data is an out‐of‐print dialect atlas (Kettunen, 1940), which we are now putting into machine‐readable form. To ease the chore of data entry, we have developed custom‐made PC‐based programs. Through a process of redundant data entry, we plan to minimize keying errors, and produce a statistically sound estimate of the number of remaining errors. Other issues we consider include: translation of the elaborate typography, generality of the data formats, intellectual property rights, and availability of original data sources. We hope that this work will provide a base not only for our work but also for other quantitative studies and computer‐based applications of this body of data.