Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding Prototypes

Xinzhou Xu, Jun Deng, Nicholas Cummins, Zixing Zhang, Li Zhao, Bjorn W. Schuller

Research output: Contribution to journalArticlepeer-review

18 Scopus citations

Abstract

Speech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes.

Original languageEnglish
Pages (from-to)2752-2765
Number of pages14
JournalIEEE Transactions on Multimedia
Volume24
DOIs
StatePublished - 2022
Externally publishedYes

Keywords

  • Speech emotion recognition
  • paralinguistics
  • semantic-embedding prototypes
  • zero-shot learning

Fingerprint

Dive into the research topics of 'Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding Prototypes'. Together they form a unique fingerprint.

Cite this