Tagging and parsing a large corpus

Svavar Kjarrval Lúthersson · 2011

This report is a product of a research where we tried to use existing language processing tools on a larger collection of Icelandic sentences than they had faced before. We hit many barriers on the way due to software errors, limitations in the software and due to the corpus we worked with. Unfortunately we had to resort to sidestep some of the problems with hacks but it resulted in a large collection of tagged and parsed sentences. We also managed to produce information regarding the frequency of words which could enhance the precision of current language processing tools.; Þessi skýrsla er afurð rannsoknar þar sem reynt er að beita nuverandi maltaeknitolum a staerra safn af islenskum setningum en aður hefur verið farið ut i. Við rakumst a ýmsar hindranir a leiðinni vegna hugbunaðarvillna, takmarkana i hugbunaðnum og vegna safnsins sem við unnum með. Þvi miður þurftum við að sneiða hja vandamalunum með ýmsum krokaleiðum en það leiddi til þess að nu er tilbuið stort safn af morkuðum og þattum setningum. Einnig sofnuðum við upplýsingum um tiðni orða sem gaetu baett nakvaemni maltaeknitola.

Read the paper · More papers on PaperTik