Huge Parsed Corpora in LASSY
Gertjan van Noord, Frank Van Eynde, Anette Frank, Koenraad De Smedt, Gertjan van Noord · 2008
One of the goals of the LASSY STEVIN project (Large Scale Syntactic Annotation of written Dutch) is a syntactically annotated (manually verified) corpus of 1 million words. In addition, the full STEVIN reference corpus of 500 million words will be syntactically annotated automatically. In this paper, the potential of such huge treebanks for applications in corpus linguistics, natural language processing and information extraction is illustrated. 1