Memory Attention: Robust Alignment Using Gating Mechanism for End-to-End Speech Synthesis
Recent end-to-end (e2e) speech synthesis systems usually employ attention techniques to align an input text sequence against a mel-spectrogram sequence. Attention-based e2e approach has shown state-of-the-art performance in speech synthesis. However, generating stable and robust attention alignment...
Saved in:
Published in | IEEE signal processing letters Vol. 27; pp. 2004 - 2008 |
---|---|
Main Authors | , , , |
Format | Journal Article |
Language | English |
Published |
New York
IEEE
2020
The Institute of Electrical and Electronics Engineers, Inc. (IEEE) |
Subjects | |
Online Access | Get full text |
Cover
Loading…
Summary: | Recent end-to-end (e2e) speech synthesis systems usually employ attention techniques to align an input text sequence against a mel-spectrogram sequence. Attention-based e2e approach has shown state-of-the-art performance in speech synthesis. However, generating stable and robust attention alignment to avoid some serious failures such as repeating, missing, and mumbling phones is still an ongoing challenge. In order to mitigate these alignment failures, we propose a novel attention method called memory attention for e2e speech synthesis, which is inspired by the gating mechanism of the long-short term memory (LSTM). Leveraging the sequence modeling power of the gating techniques, memory attention can produce a stable alignment by controlling the amount of content-based and location-based information. For performance evaluation, we compared our proposed memory attention algorithm with various conventional attention techniques in single speaker and emotional speech synthesis scenarios. From the experimental results, we conclude that memory attention can robustly generate various stylish speech. |
---|---|
ISSN: | 1070-9908 1558-2361 |
DOI: | 10.1109/LSP.2020.3036349 |