The LDC-IL Speech Corpora
Narayan Kumar Choudhary, Durgesh Rao · 2020
This paper introduces the first set of speech corpora released in 2019 by the Linguistic Data Consortium for Indian Languages (LDC-IL), a scheme under the Department of Higher Education, Ministry of Human Resource Development, Government of India. The datasets include a total of 13 scheduled languages of India, collected in various environments across length and breadth of the vast country, from a total of 5662 speakers of different age-groups with a total size of more than 1552 hours. The dataset is still growing as we prune them and make them ready for release. Unique language corpus is usually the largest available at present for these languages. Established in 2008, on the lines of the LDC of University of Pennsylvania, the LDC-IL has worked for over 10 years on various types language resources, including building the speech corpora. LDC-IL is a fully government funded project implemented by CIIL, Mysuru. Due to some restraints in the government business such as cost analysis and copyright issues, it took rather a long time to release the LDC-IL dataset for the public use. This paper gives a brief of the raw speech corpora now released and ready for public use (both commercial and non-commercial purposes). It also discusses how the two major bottlenecks of copyright and costing was addressed which held up the release of these datasets for several years.