Recognition of Audio Depression Based on Convolutional Neural Network and Generative Antagonism Network Model

This paper proposes an audio depression recognition method based on convolution neural network and generative antagonism network model. First of all, preprocess the data set, remove the long-term mute segments in the data set, and splice the rest into a new audio file. Then, the features of speech s...

Full description

Saved in:

Bibliographic Details
Published in	IEEE access Vol. 8; pp. 101181 - 101191
Main Authors	Wang, Zhiyong, Chen, Longxi, Wang, Lifeng, Diao, Guangqiang
Format	Journal Article
Language	English
Published	Piscataway IEEE 2020 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subjects	Algorithms Artificial neural networks Audio data Convolution convolutional neural network Data models Datasets Entropy entropy feature of spectrogram Feature extraction Filter banks generative antagonism network Hidden Markov models Machine learning Matrix algebra Matrix methods Mel-scale frequency cepstral coefficients Neural networks Recognition Recognition of audio depression Root-mean-square errors Segments Speech recognition
Online Access	Get full text

Cover

Loading…

More Information
Summary:	This paper proposes an audio depression recognition method based on convolution neural network and generative antagonism network model. First of all, preprocess the data set, remove the long-term mute segments in the data set, and splice the rest into a new audio file. Then, the features of speech signal, such as Mel-scale Frequency Cepstral Coefficients (MFCCs), short-term energy and spectral entropy, are extracted based on audio difference normalization algorithm. The extracted matrix vector feature data, which represents the unique attributes of the subjects' own voice, is the data base for model training. Then, based on the combination of CNN and GAN, DR AudioNet is used to build the model of depression recognition research. With the help of DR AudioNet, the former model is optimized and the recognition classification is completed through the normalization characteristics of the two adjacent segments before and after the current audio segment. The experimental results on AViD-Corpus and DAIC-WOZ datasets show that the proposed method effectively reduces the depression recognition error compared with other existing methods, and the RMSE and MAE values obtained on the two datasets are better than the comparison algorithm by more than 5%.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14
ISSN:	2169-3536 2169-3536
DOI:	10.1109/ACCESS.2020.2998532