From digital library to n-grams: NB N-gram
Magnus Breder Birkenes, Lars G. Johnsen, Arne Martinus Lindstad, Johanne Ostad · 2015
At the National Library of Norway, we are currently developing a service comparable to the Google Ngram Viewer (Michel et al., 2010; Lin et al., 2012; Aiden and Michel, 2013) called NB Ngram. It is based on all books and newspapers digitized up to and including 2013, as part of the large scale digitization project at the National Library of Norway. Uni-, bi- and trigams have been generated on the basis of this text corpus containing some 34 billion words. In this paper, we sketch the background of NB N-gram and illustrate some applications of it. 1 Background In 2006, the National Library of Norway initiated an ambitious digitization program, with the goal of digitizing its entire collection. The collection contains all material collected under the legal deposit act, and includes among other things books, newspapers, periodicals, magazines, journals, music, films, posters and maps; basically anything published in the public domain in more than 50 copies. The collection contains material in many different languages.