vix.ing · top · new · best · stats · spec

Development and suitability of Indian languages speech database for building watson based ASR system

2013/11/01 by Dipti Pandey, Tapabrata Mondal, S. S. Agrawal +1 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques

paper · doi:10.1109/icsda.2013.6709861

openalex publication_date 2013/11/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

In this paper, we discuss our efforts in the development of Indian spoken languages corpora for building large vocabulary speech recognition systems using WATSON Toolkit. The current paper demonstrates that these corpora can be reduced to a varied degree for various phonemes by comparing the similarity among phonemes of different languages. We also discuss the design and methodology of collection of speech databases and the challenges we have faced during database creation. The experiments have been conducted on commonly known Indian languages, by training the ASR system with WATSON toolkit and evaluation by Sclite. The results for these experiments show that different Indian languages have a great similarity among their phoneme structures and phoneme sequences and we have exploited these features to create speech recognition system. Also, we have developed an algorithm to bootstrapping the phonemes of one particular language into another by mapping the phonemes of different languages. The performance of Hindi and Bangla ASR systems using these databases has been compared.

Cited by

Related