Finite State Models for the Generation of Large Corpora of Natural Language Texts
Domenico Aldo Cantone, Salvatore Cristofaro, Simone Faro, Emanuele Giaquinta · Frontiers in artificial intelligence and applications · 2009
Natural languages are probably one of the most common type of input for text processing algorithms. Therefore, it is often desirable to have a large training/testing set of input of this kind, especially when dealing with algorithms tuned for natural language texts. In many cases the problem due to the lack of big corpus of natural language texts can be solved by simply concatenating a set of collected texts, even with heterogeneous contexts and by different authors.