BenchmarkDataNLP.jl: Synthetic Data Generation for NLP Benchmarking
Alexander V. Mantzaris · The Journal of Open Source Software · 2025
BenchmarkDataNLP.jl is a package written in Julia Lang for generating synthetic text corpora that can be used to systematically benchmark and evaluate Natural Language Processing (NLP) models such as Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTMs), and Large Language Models (LLMs).By enabling users to control core linguistic parameters such as the alphabet size (characters selected from the Unicode Hangul block, Korean Language), vocabulary size, grammatical expansion complexity, and semantic structures this library can help users test, evaluate, and debug NLP models.Instead of exposing many parameters to the user, it is kept at minimum, so that users do not have to focus on how to correctly configure the generation process.The key parameter is the complexity (integer value), which controls the size of the alphabet, vocabulary, and grammar expansions.This parameter accepts an integer in the range 1 to 100.For example, at complexity = 1 (simplest value), there are 5 letters in the alphabet and 10 words used in the vocabulary with 2 grammar roles when a Context Free Grammar generator is selected.At complexity = 100 there is a vocabulary of 10,000 words and 50 alphabet characters.Users can choose how many independent grammar productions are desired, which are each supplied as entries in a .jsonlfile.The defaults provided in the documentation should suffice for most use cases.