Monophone and Triphone Acoustic Phonetic Model for Kannada Speech Recognition System a Mahesh Kumar T N,b Adithya Jayan,b Shreenidhi Bhat,b Anvith M,c A V Narasimhadhan Department of Electronics and Communication Engineering National Institute of Technology, Surathkal.Karnataka ,India. a,b (mahesh.ssy100,adithyaj8,shribhat27,anvithmkulal2000)@gmail.com,c dhansiva@nitk.edu.in Abstract—The automatic Speech Recognition system (ASR) is the most widely used application in the speech domain. ASR systems generate text data from spoken utterances without manual intervention. In this work, we build an ASR system for the Kannada language. For building the proposed system, we extract Mel Frequency Cepstral Coefficients (MFCC) features from the audio data, and the Kannada language model is developed using corresponding labels. The dictionary generation and phonetic labelings are automated. Recognition performance is compared for both monophonic and triphone models. The word error rate of 15.73% and the sentence error rate of 55.5% are achieved for the triphone model. Comparatively, the triphone model gives a better performance than the monophonic model. Index Terms—Automatic Speech Recognition, MFCC, HMM, Kannada language I. I NTRODUCTION The objective of the Automatic Speech Recognition system is to recognize the speech and map it to the corresponding text. It mainly involves converting audio to text irrespective of the speaking style, accent, and other characteristics. Among different approaches to speech recognition, the Acoustic-Phonetic approach [1] segments the speech into stable acoustic units and, based on the acoustic properties of each segment, phonetic labels are attached to it. This method requires breaking the word strings into specific phone units. Hence the model must be able to extract acoustic characteristics efficiently. With the increased research in the Speech Recognition domain, the development of large vocabulary systems is easier by using phonemes as a basic unit of acoustic properties. Most languages can be represented efficiently using a finite number of phonemes. This makes phonemes an obvious choice while building ASR systems. Most languages show some amount of correlation between the order of the occurrences of phones. Thus grouping of phones can prove to be advantageous. Triphone modeling [2] is one such grouping, where each phone is grouped with the previous and the following phones to improve performance. Even with a finite number of phones, the number of possible triphones that can be formed is enormous. Thus we resort to clustering to obtain a reasonably good recognition rate. Many tools and software are available for speech recognition and ASR development, such as The Hidden Markov Model toolkit [3], Julius [4], Sphinx [5] and The Kaldi ASR toolkit [6]. Kaldi is an open-source toolkit for Speech Recognition written in C++. It is a robust toolkit that offers 978-1-6654-9648-3/22/$31.00 © 2022 IEEE an FST (Finite state transducer) based framework, extensive linear algebra support, and also a non-restrictive license. The Kaldi algorithm is presented in a generic format that allows flexibility while designing and building the model, and hence the proposed model is built using Kaldi. The paper is organized as follows: We started with a literature survey, Proposed Kannada ASR System which includes Kannada language phonology, MFCC Feature extraction, Acoustic Model and Language Model, Implementation part includes data set used, i.e., speech corpus, Performance Metrics, Results and Discussion,and Conclusion. II. L ITERATURE REVIEW Automatic speech recognition is a domain is of great potential. It was not until the 1980s where there was a breakthrough in statistics that introduced Hidden Markov Models (HMMs) [7], [8] that changed the way we looked at speech signals is more mathematical and probabilistic models. Currently, there exist various models, and toolkits such as HTK, Julius, Sphinx4, RWTH ASR Toolkits [9] for automatic speech recognition. The Kaldi toolkit is one such toolkit that has efficient opensource models for natural language processing which is under continuous development. Multiple studies have been carried out to recognize languages like Punjabi, Hindi, Tamil, Telugu, Kannada, and others. [10] [11]. Some work related to speech recognition in Hindi using Kaldi is reported in [12]. A popular method for ASR is the acoustic-phonetic method. The phonetic labels are related to acoustic segments based on features in this method. The phonetic labels are then mapped to words to recognize the speech signal. Multiple papers have explored this method [10], [11]. The third method that is currently gaining popularity is using deep learning models [13]. Extensive research has been done on ASR for the English language. However, native Indian Languages like Kannada, due to the less availability of the data and other reasons, have not received the same attention and have a great scope of improvement. Though to some extent, slight dialectal variations do not much affect the ASR system, but more variations in the dialect may affect the ASR results [14]. Research published for speech recognition on Indian languages such as Hindi, Telugu, Malayalam, and Kannada have shown acceptable results [11], [12]. A paper on Continuous Speech Recognition for the Kannada Language [10] has achieved good results using the Triphone modeling technique. In this paper, we discuss the data preparation and preprocessing for a Kannada language dataset. Monophone and triphone-based ASR models for recognition of the Kannada have been presented, and their performance has been evaluated and compared. III. P ROPOSED K ANNADA ASR SYSTEM In this section, the details of the developed Kannada ASR system is given. This section includes the brief details of stages involved in the development of ASR system. A block diagram of the ASR system flow is given in Fig.1. B. MFCC Extraction Mel-frequency cepstral coefficients (MFCC) is a popular audio feature extraction method, where the coefficients are extracted from the frequency components of the speech signal at specific bands. The general MFCC extraction flow is described in Fig.4. To extract audio features from an audio segment, we use a sliding window of 25ms width, and the features are extracted from each slide(10ms). The frequency spectrum of these audio windows is then extracted. The spectrum of each audio window is passed through a series of band-pass filters. Fig.5 shows the Kaldi implementation of feature extraction. Fig. 4. Mel Frequency Cepstral Coefficients extraction flowchart Fig. 1. ASR Block diagram A. Kannada language phonology The Kannada language consists of 12 vowels (Swaras), 2 diphthongs, 2 yogavahakas. The Consonant set consists of 25 vargiya vyanjana (Categorized consonant) and 11 avargiya vyanjana (Sonorants). So totally in 52 phones are found in the Kannada language. Out of this, in modern Kannada, three are obsolete and marked in Fig.2. Unlike English, most Indian languages have a one-toone correlation between sounds and letters, i.e., the Indian languages scripts are phonetic. Thus the phonetic dictionary construction from the transcription can be automated to a great extent. A list of letters and the associated phonetic labels used for them is shown in Fig.3 The phones used were according to the guidelines specified by the ILSL (Indian Language Speech sound Label set) version 3.1 guidelines [15]. It provides a standard for phonetic labeling of speech sounds for various Indian languages. Fig. 5. Feature Extraction Module Flowchart C. Acoustic model The acoustic model is used in speech recognition to map the audio recordings with the corresponding phoneme representation. The model is trained on training audio and its corresponding text corpus. The model represents each phoneme using HMM parameters, which are a representation of the Gaussian distribution for the corresponding phoneme. The audio signal is mapped to that phoneme whose Gaussian distribution is nearest to the HMM parameters for the corresponding audio signal in the vector space. The HMM parameters from the Acoustic model and the probability distribution from the Language model lets us find the optimal word sequence as shown below: Let W ∗ be the optimal word sequence from the vocabulary, X be the input sequence of acoustic features and W be the word sequence in language model. Fig. 2. Table of Kannada Letters W ∗ = argmaxW P (W |X) (1) Fig. 3. Table mapping Kannada Letters to their Phonetic Representation Equation (1) is known as ”Fundamental Equation of Statistical Speech Processing”. Applying Bayes theorem, we get (2). W ∗ = argmaxW P (X|W )P (W ) P (X) (2) We can re-formulate (2) as shown in (3). W ∗ = argmaxW P (X|W )P (W ) the model learns the dependencies between phonemes/words in the given text data. We believe our audio clips are grammatically and semantically correct, and as a result, incorporating a language model into decoding will increase ASR accuracy. The probability of a sequence of words is calculated using the language model. The next word is predicted using the sequence of previous ’n’ words. This is known as n-gram modelling. (3) In equations (1), (2), (3) P (X|W ) represents the probability that feature vector X maps to the word W (acoustic model) and P (W ) represents the probability of the word W occurring next in sequence(language model). 1) Monophone modelling: On an average, a phone is uttered by a speaker every 10-30 ms. Hence, audio is sampled every 10 ms with a window size of 25 ms, such that the sampled audio signal represents a phoneme. This is followed by MFCC feature extraction, followed by calculating the HMM parameters. The appropriate phoneme is mapped based on these extracted HMM parameters. 2) Triphone modelling: There are chances that a phoneme may not be present in the central part of the windowed audio signal. Hence, the triphone modelling takes into account the MFCC feature vectors extracted from the previous windowed audio and the next windowed audio signal by concatenating them with the MFCC feature vectors for current windowed audio. This is followed by calculating the HMM parameters and mapping the corresponding phoneme. D. Language model synthesis The language model provides a prediction of the next sequence of phones based on the current and/or previous set of phones. The model is trained with the corpus text, therefore IV. I MPLEMENTATION Kaldi is a robust and well developed open-source speech recognition module that supports context-dependent phone modelling (Acoustic-Phonetic Modeling). It processes the speech data to get the MFCC coefficients which are trained and then fitted with Language Modeling for mapping acoustic results to the Phones. Kaldi also supports Subspace Gaussian Mixture Model (SGMM) systems. A. Speech Corpus The speech data were obtained from Google research Datasets published under the CC BY-SA 4.0 attribution [16]. The audio is sampled at 48000 samples per second with 16 bits per sample. The corpus consisted of high-quality transcribed recordings of approximately 4,400 sentences which included both male and female voices in equal proportions. Corresponding transcriptions were provided in the Kannada language. The entire set was divided into test and train sets, with 100 recordings taken from both the male and female sets to use as the test set. The sampling of 48000 per second is down sampled to 16000 samples per second. Similarly quantization of 16 bit is made to 8 bit. The dataset transcriptions are cleaned by removing punctuation and annotations. Numbers were either replaced with their equivalent in text or removed entirely from the dataset. This is followed by creating a phonetic translation of the corpus according to the ILSL guidelines. The generated transcriptions are transformed into the required formats for training. The word dictionary, word-level phonetic transcription, text corpus, and files necessary to map transcripts to appropriate audio files are generated. Scripts from the Kaldi toolkit were used for format verification and data sorting. The generated dictionary consisted of a total of 10451 different words and their phonetic translations. A total of 50 non-silence phones were used, along with two silence phones. One phoneme for the short silence (sil) was used as the default for unknown sounds, and (en) was used for ligature (othhakshara). The final processed corpus consisted of 4003 sentences and their utterances for training and 200 for testing. B. Performance Metrics As metrics, the final predictions are evaluated using WER (word error rate) and SER (Sentence error rate). The definition of WER is given in (4) W ordErrorRate = (S + I + D)/N (4) In Equation (4), S (Substitutions) is the number of words that were mispredicted as a different word. I (Insertions) is the number of additional words not present in the sentence that was added, and D (Deletions) is the number of words wholly missed out in the prediction. N is the total number of words in the utterance + Insertions. Thus, WER acts as a good metric for performance evaluation, where a lower WER implies a better prediction model. SER stands for Sentence error rate. It is defined as SER = 1 − n(totallycorrectutt)/n(utt) (5) The SER is essentially the percentage of predictions that are not completely perfect. This means that SER considers that utterance entirely wrong even if a single word is mispredicted. Due to this, the SER is generally a lot higher than the WER, and a lower SER implies better sentence predictions. Fig. 6. Performance comparison of the SER and WER TABLE I R ESULTS AND C OMPARISON Model\Performance Monophone Triphone WER(%) 28.58 15.73 SER(%) 72.0 55.5 We observe that the triphone model performs significantly better than the monophone model with a WER of 15.73% and an SER of 55.5%. The performance can be attributed to the triphone model’s better context-dependent phone prediction. VI. C ONCLUSION In this paper, an automatic speech recognition system is implemented using the Kaldi toolkit. The extracted MFCC features are used to create monophonic and triphone models. Good performance is observed for both models. Out of two developed models, triphone based model showed better performance. The performance may be improved by using more data and developing state-of-the-art deep neural network based systems. VII. ACKNOWLEDGEMENTS The authors would like to thank Sathvik Bhat and Thilak Shetty for their contributions to the project. We also thank National Institute of Technology Karnataka, for providing the opportunity and resources to conduct this project. V. R ESULTS AND D ISCUSSION R EFERENCES The test data undergo a similar procedure as the train data. The MFCC features are extracted first. The decode graph that is generated using the training data is then used to decode the test data. These predictions are then evaluated to judge the model efficiency. The performance comparison of the SER and WER between monophone and triphone models is shown in Fig.6 The results obtained are tabulated in the table I ,and shown graphically in Fig.6. The monophone model obtained a WER of 28.58% and an SER of 72.0%. The triphone model obtained a WER 15.73% of and an SER of 55.5%. [1] R. Schwartz, Y. Chow, O. Kimball, S. Roucos, M. Krasner, and J. Makhoul, “Context-dependent modeling for acoustic-phonetic recognition of continuous speech,” in ICASSP’85. IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 10, pp. 1205– 1208, IEEE, 1985. [2] S. Dupont, C. Ris, L. Couvreur, and J.-M. Boite, “A study of implicit and explicit modeling of coarticulation and pronunciation variation,” in Ninth European Conference on Speech Communication and Technology, 2005. [3] S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey, et al., “The htk book,” Cambridge university engineering department, vol. 3, no. 175, p. 12, 2002. [4] A. Lee, T. Kawahara, and K. Shikano, “Julius—an open source real-time large vocabulary recognition engine,” 2001. [5] W. Walker, P. Lamere, P. Kwok, B. Raj, R. Singh, E. Gouvea, P. Wolf, and J. Woelfel, “Sphinx-4: A flexible open source framework for speech recognition,” 2004. [6] D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al., “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding, no. CONF, IEEE Signal Processing Society, 2011. [7] L. Rabiner and B. Juang, “An introduction to hidden markov models,” ieee assp magazine, vol. 3, no. 1, pp. 4–16, 1986. [8] K. Riedhammer, T. Bocklet, A. Ghoshal, and D. Povey, “Revisiting semi-continuous hidden markov models,” in 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4721–4724, IEEE, 2012. [9] D. Rybach, C. Gollan, G. Heigold, B. Hoffmeister, J. Lööf, R. Schlüter, and H. Ney, “The rwth aachen university open source speech recognition system,” in Tenth Annual Conference of the International Speech Communication Association, 2009. [10] S. C. Sajjan and C. Vijaya, “Continuous speech recognition of kannada language using triphone modeling,” in 2016 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET), pp. 451–455, IEEE, 2016. [11] L. B. Babu, A. George, K. Sreelakshmi, and L. Mary, “Continuous speech recognition system for malayalam language using kaldi,” in 2018 International Conference on Emerging Trends and Innovations In Engineering And Technological Research (ICETIETR), pp. 1–4, IEEE, 2018. [12] P. Upadhyaya, O. Farooq, M. R. Abidi, and Y. V. Varshney, “Continuous hindi speech recognition model based on kaldi asr toolkit,” in 2017 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET), pp. 786–789, IEEE, 2017. [13] Z. Zhang, J. Geiger, J. Pohjalainen, A. E.-D. Mousa, W. Jin, and B. Schuller, “Deep learning for environmentally robust speech recognition: An overview of recent developments,” ACM Transactions on Intelligent Systems and Technology (TIST), vol. 9, no. 5, pp. 1–28, 2018. [14] P. Hegde, N. B. Chittaragi, S. K. P. Mothukuri, and S. G. Koolagudi, “Kannada dialect classification using cnn,” in International Conference on Mining Intelligence and Knowledge Exploration, pp. 254–259, Springer, 2019. [15] K. Samudravijaya, “Indian language speech label (ilsl): a de facto national standard,” in Advances in Speech and Music Technology, pp. 449–460, Springer, 2021. [16] F. He, S.-H. C. Chu, O. Kjartansson, C. Rivera, A. Katanova, A. Gutkin, I. Demirsahin, C. Johny, M. Jansche, S. Sarin, et al., “Open-source multi-speaker speech corpora for building gujarati, kannada, malayalam, marathi, tamil and telugu speech synthesis systems,” in Proceedings of the 12th Language Resources and Evaluation Conference, pp. 6494– 6503, 2020.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )