JOURNAL OF CHINA UNIVERSITIES OF POSTS AND TELECOM ›› 2017, Vol. 24 ›› Issue (2): 1-9.doi: 10.1016/S1005-8885(17)60193-6

    Next Articles

Noisy speech emotion recognition using sample reconstruction and multiple-kernel learning

  

  • Received:2016-09-29 Revised:2017-04-01 Online:2017-04-30 Published:2017-04-30
  • Supported by:
    the National Natural Science Foundation of China (61501204, 61601198), the Hebei Province Natural Science Foundation (E2016202341), the Hebei Province Foundation for Returned Scholars (C2012003038), the Shandong Province Natural Science Foundation (ZR2015FL010), the Science and Technology Program of University of Jinan (XKY1710).

Abstract: Speech emotion recognition (SER) in noisy environment is a vital issue in artificial intelligence (AI). In this paper, the reconstruction of speech samples removes the added noise. Acoustic features extracted from the reconstructed samples are selected to build an optimal feature subset with better emotional recognizability. A multiple-kernel (MK) support vector machine (SVM) classifier solved by semi-definite programming (SDP) is adopted in SER procedure. The proposed method in this paper is demonstrated on Berlin Database of Emotional Speech. Recognition accuracies of the original, noisy, and reconstructed samples classified by both single-kernel (SK) and MK classifiers are compared and analyzed. The experimental results show that the proposed method is effective and robust when noise exists.

Key words: speech emotion recognition, compressed sensing, multiple-kernel learning, feature selection

CLC Number: