Pashto language dialect recognition using mel frequency cepstral coefficient and support vector machines

Saud Khan, Haider Ali, Khalil Ullah · 2017

In this paper, a system based on support vector machines is proposed for content-based dialect classification and retrieval. This work is part of an ongoing effort to address the needs of new under-resourced languages. The recognition system will work for the interest and welfare of the Pashto speaking people and will help in keeping the language dialects alive by this process. Voice samples are collected for the purpose of making a database which includes dialects from different age groups and genders. The feature sets extracted from the database includes cepstral coefficients and statistical parameters which are defined by optimum class boundaries using support vector machines. Experiments show that a support vector machines based dialect recognition system gives optimum result in accurately distinguishing between different dialects.

Read the paper · More papers on PaperTik