Speech recognition of mandarin monosyllables

作者：

Highlights：

•

摘要

The nonlinear dynamic characteristics of expansion and contraction and the sequential time-varying features of the syllable pronunciations greatly complicate the tasks of automatic speech recognition. Each syllable is represented by a sequence of vectors of linear predict coding cepstra (LPCC). Even if the same speaker utters the same syllable, the duration of stable parts of the sequence of LPCC vectors changes every time. Therefore, the duration of stable parts is contracted such that the compressed speech waveform has the same length. We propose five different simple techniques to contract the stable parts of the sequence of LPCC vectors. A simplified Bayes decision rule with a weighted variance is used to classify 408 speaker-dependent mandarin syllables. For the 408 speaker-dependent mandarin syllables, the recognition rate is 94.36% as compared with 79.78% obtained by using the hidden Markov models (HMM). A recognition rate 98.16% is achieved within top 3 candidates. The features proposed in this paper to represent the syllables are simple and easy to be extracted. The computation for feature extraction and classification is much faster than using the techniques of the HMM or any other known techniques.

论文关键词：Bayes decision rule,Linear predict coding,Speech recognition

论文评审过程：Received 24 June 2002, Revised 13 March 2003, Accepted 4 April 2003, Available online 27 June 2003.

论文官网地址：https://doi.org/10.1016/S0031-3203(03)00135-3