Alternating Synthetic and Real Gradients for Neural Language Modeling

Training recurrent neural networks (RNNs) with backpropagation through time (BPTT) has known drawbacks such as being difficult to capture longterm dependencies in sequences. Successful alternatives to BPTT have not yet been discovered. Recently, BP with synthetic gradients by a decoupled neural inte...

Full description

Saved in:
Bibliographic Details
Published inarXiv.org
Main Authors Shang, Fangxin, Zhang, Hao
Format Paper
LanguageEnglish
Published Ithaca Cornell University Library, arXiv.org 03.06.2022
Subjects
Online AccessGet full text

Cover

Loading…