Textual Datasets For Portuguese-Brazilian Language Models

Matheus Ferraroni Sanches, Jader Martins Camboim de Sá, Henrique Theodor Schutz Foerste, Rafael Roque de Souza, Júlio Cesar dos Reis, Leandro Aparecido Villas · 2022

Advances in Natural Language Processing have generated new models that push forward the state of the art. This reached new heights in complex tasks in handling unstructured texts. Most of the new architectures and models focus on the English language. There is a lack of available datasets that can be used during the training of new models. This investigation presents four new textual datasets for language modeling in Brazilian Portuguese. Our datasets were generated from several specific methodologies that aimed to obtain data of different natures. Two of our sets were originally built from data in online web forums. We also distribute a translated version of MultiWOZ, and a clean version of BrWaC. The original datasets are made available in a structured way to facilitate their use during the training of NLP models, with questions, answers and conversations already identified.

Read the paper · More papers on PaperTik