An automatic analysis of Norwegian compounds
Janne Bondi Johannessen, Helge Hauglin · NORA - Norwegian Open Research Archives · 1998
Introduction The University of Oslo is currently developing an automatic morphosyntactic tagger for Norwegian. 2 A very important module is one which can analyse compounds. Compounding is extremely productive in Norwegian, and it is futile to ever hope for a lexicon (dictionary) that will contain all or even most of the compounds that occr in actual texts. Since the tagger we are developing is based on the possibility of recognising words by the help of a lexicon, it is of great importance to have a module that recognises new compounds. According to Munthe (1972), 10.4 per cent of all words in running text are compounds. Any text sample will contain a greatnumber of compounds. This statistics is true even for small samples. I took an arbitrary 440-word article from the newspaper Aftenposten from September this year, Full penhet om passunion (Full openness on passport union), and I quickly counted 47 compounds. Many of them a