VLibrasBD: A Brazilian Portuguese–Brazilian sign language (Libras) bilingual text dataset designed to support neural machine translation

Manuella Aschoff Cavalcanti Brandão Lima, Daniel Cruz, Diego R. B. da Silva, Dilainne Daniel Albuquerque, Daniel Faustino Lacerda, Rostand Edson Oliveira Costa, Guido Lemos de Souza Filho, Tiago Maritan Ugulino de Araújo · Data in Brief · 2025

VLibras-DB is a bilingual text corpus in Brazilian Portuguese (BP) and Brazilian Sign Language (Libras), designed and developed to support the creation of machine translation systems from BP-to-Libras. The corpus adopts a textual notation for Libras known as gloss, which serves as an interlingua between the source and target languages. To support this process, we initially defined a set of grammatical rules specific to Libras. Based on this notation, a bilingual textual database was built by a team of ten Libras interpreters, resulting in a corpus comprising 127,349 BP-Libras translation pairs. The dataset includes approximately 72,000 general-purpose sentences and around 55,000 sentences extracted from Brazilian federal government content and services.. The dataset was carefully constructed to include a wide variety of lexical and syntactic phenomena relevant to Libras translation, such as directional verbs, intensifiers, negation, and word-sense disambiguation. The resulting resource provides not only a substantial volume of parallel data but also a linguistically informed foundation for training and evaluating NMT models, contributing significantly to the advancement of accessible language technologies for the Deaf community. This comprehensive dataset is particularly significant for Neural Machine Translation (NMT) as it provides a much-needed, high-quality resource to train and evaluate NMT models for this low-resource language pair, facilitating advancements in BP-to-Libras translation systems. Beyond its direct application in NMT, VLibrasBD serves as a foundational linguistic resource for natural language processing, supporting tasks such as comparative linguistic analysis, bilingual embedding training, and the development of assistive technologies to enhance multilingual communication and information accessibility.

Read the paper · More papers on PaperTik