Deep CNNs Along the Time Axis With Intermap Pooling for Robustness to Spectral Variations

Convolutional neural networks (CNNs) with convolutional and pooling operations along the frequency axis have been proposed to attain invariance to frequency shifts of features. However, this is inappropriate with regard to the fact that acoustic features vary in frequency. In this paper, we contend...

Full description

Saved in:

Bibliographic Details
Published in	IEEE signal processing letters Vol. 23; no. 10; pp. 1310 - 1314
Main Authors	Lee, Hwaran, Kim, Geonmin, Kim, Ho-Gyeong, Oh, Sang-Hoon, Lee, Soo-Young
Format	Journal Article
Language	English
Published	New York IEEE 01.10.2016 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects	Acoustic modeling Acoustics Architecture Convolution convolutional neural networks (CNNs) Feature maps Frequency shift Hidden Markov models IMP intermap pooling (IMP) layer Neural networks Robustness Spectra Speech Tasks Training
Online Access	Get full text

Cover

Loading…

More Information
Summary:	Convolutional neural networks (CNNs) with convolutional and pooling operations along the frequency axis have been proposed to attain invariance to frequency shifts of features. However, this is inappropriate with regard to the fact that acoustic features vary in frequency. In this paper, we contend that convolution along the time axis is more effective. We also propose the addition of an intermap pooling (IMP) layer to deep CNNs. In this layer, filters in each group extract common but spectrally variant features, then the layer pools the feature maps of each group. As a result, the proposed IMP CNN can achieve insensitivity to spectral variations characteristic of different speakers and utterances. The effectiveness of the IMP CNN architecture is demonstrated on several LVCSR tasks. Even without speaker adaptation techniques, the architecture achieved a WER of 12.7% on the SWB part of the Hub5'2000 evaluation test set, which is competitive with other state-of-the-art methods.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 23
ISSN:	1070-9908 1558-2361
DOI:	10.1109/LSP.2016.2589962