Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

Jing Han, Zixing Zhang, Zhao Ren, Bjorn Schuller

Publikation: Beitrag in Buch/Bericht/KonferenzbandKonferenzbeitragBegutachtung

29 Zitate (Scopus)

Abstract

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the training procedure for either speech or facial emotion recognition. Specifically, the model consists of one modality-specific network per individual modality and one shared network to map both audio and visual cues into final predictions. In the training process, we additionally take the loss from one auxiliary modality into account besides the main modality. To evaluate the effectiveness of the implicit fusion model, we conduct extensive experiments for mono-modal emotion classification and regression, and find that the implicit fusion models outperform the standard mono-modal training process.

OriginalspracheEnglisch
Titel2019 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Proceedings
Herausgeber (Verlag)Institute of Electrical and Electronics Engineers Inc.
Seiten5861-5865
Seitenumfang5
ISBN (elektronisch)9781479981311
DOIs
PublikationsstatusVeröffentlicht - Mai 2019
Extern publiziertJa
Veranstaltung44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019 - Brighton, Großbritannien/Vereinigtes Königreich
Dauer: 12 Mai 201917 Mai 2019

Publikationsreihe

NameICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Band2019-May
ISSN (Print)1520-6149

Konferenz

Konferenz44th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2019
Land/GebietGroßbritannien/Vereinigtes Königreich
OrtBrighton
Zeitraum12/05/1917/05/19

Fingerprint

Untersuchen Sie die Forschungsthemen von „Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality“. Zusammen bilden sie einen einzigartigen Fingerprint.

Dieses zitieren