Rank-frequency distributions of Romanian words
Adrian Cocioceanu, Mihaela Carina Raportaru, Alexandru I. Nicolin, Dragan Jakimovski · AIP conference proceedings · 2017
The calibration of voice biometrics solutions requires detailed analyses of spoken texts and in this context we investigate by computational means the rank-frequency distributions of Romanian words and word series to determine the most common words and word series of the language. To this end, we have constructed a corpus of approximately 2.5 million words and then determined that the rank-frequency distributions of the Romanian words, as well as series of two, and three subsequent words, obey the celebrated Zipf law.