Natural Language Processing Methods for Language Modeling

Dávid Márk Nemeskey · Eötvös Loránd Tudományegyetem · 2021

In this thesis we concentrated on two issues: how modern NLP (or traditional linguistic) techniques can improve language modeling, and how to improve the state of the art for Hungarian. Chapter 2 highlighted the problems of word-based language modeling for Hungarian and introduced the “gluten-free” format, a morphological segmentation algorithm that alleviates the adverse effects of the overabundance of word forms in the language. Chapter 3 proposed a novel method for evaluating multi-sense embeddings based on lexicographical resources. Chapter 5 gave an example for the opposite direction, when language models are used to improve the performance of an NLP system. We improved on the state of the art in Hungarian language modeling in several ways. First, we presented a set of language modeling benchmarks on three Hungarian corpora in Chapter 2. A preprocessed version of the Hungarian Webcorpus has been released to serve as a standard dataset for language model assessment. A new Hungarian corpus has been created in Chapter 4. Webcorpus 2.0 was compiled from Hungarian pages in the Common Crawl and the Hungarian Wikipedia. At 9 billion tokens, it is 3.5 times the size of the previous largest (commercial) corpus, and can serve as training data for large-scale language models. Finally, the emBERT module developed in Chapter 5 enables the integration of modern contextualized embedding-based classifiers into the e-magyar pipeline. Based on our preliminary Hungarian BERT model, the NP chunker outperforms the previous best system by 2.8% in F1 score. All resources (corpora, models and software alike) presented in the thesis are freely downloadable under permissive licenses.

Read the paper · More papers on PaperTik