Discriminative dynamic Gaussian mixture selection with enhanced robustness and performance for multi-accent speech recognition

We propose a discriminative DGMS (dynamic Gaussian mixture selection) strategy to enhance restructuring of a pre-trained set of Gaussian mixture models to cover unexpected acoustic variations at run time in automatic speech recognition. The number of Gaussian components in each hidden Markov model (...

Full description

Saved in:

Bibliographic Details
Published in	2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) pp. 4749 - 4752
Main Authors	Chao Zhang, Yi Liu, Yunqing Xia, Chin-Hui Lee
Format	Conference Proceeding
Language	English
Published	IEEE 01.03.2012
Subjects	Accent Acoustics Biological cells Dynamic Gaussian Mixture Selection Genetic Algorithm Genetic algorithms Hidden Markov models Minimum Classification Error Speech Speech recognition Vectors
Online Access	Get full text

Cover

Loading…

More Information
Summary:	We propose a discriminative DGMS (dynamic Gaussian mixture selection) strategy to enhance restructuring of a pre-trained set of Gaussian mixture models to cover unexpected acoustic variations at run time in automatic speech recognition. The number of Gaussian components in each hidden Markov model (HMM) state set aside is determined by a minimum classification error criterion. We also use a genetic algorithm to solve the integer programing problem to find the globally optimal state size. This parameter is used to adjust the HMM state densities for each input speech frame, leading to both high robustness and good resolution for dynamic tracking to cover a diversity of temporal variations in speech. Tested on an accented speech recognition application, the proposed framework yields an improved syllable error rate reduction over the conventional DGMS and augmented HMM systems when evaluated on three typical Chinese accents, Chuan, Yue and Wu, while maintaining its performance for standard Putonghua.
ISBN:	1467300454 9781467300452
ISSN:	1520-6149 2379-190X
DOI:	10.1109/ICASSP.2012.6288980