Adaptive Alignment and Time Aggregation Network for Speech-Visual Emotion Recognition

IF 3.2 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC IEEE Signal Processing Letters Pub Date : 2025-03-10 DOI:10.1109/LSP.2025.3550007

Lile Wu;Lei Bai;Wenhao Cheng;Zutian Cheng;Guanghui Chen

引用次数: 0

Abstract

Video-based speech-visual emotion recognition plays a crucial role in human-computer interaction applications. However, it faces several challenges, including: 1) the redundancy in the extracted speech-visual features caused by the heterogeneity between speech and visual modalities, and 2) the ineffective modeling of the time-varying characteristics of emotions. To this end, this paper proposes an adaptive alignment and time aggregation network (AataNet). Specifically, AataNet designs a low redundancy speech-visual adaptive alignment (LRSVAA) module to acquire the low-redundant aligned features of speech-visual modalities. Meanwhile, AataNet also designs a computationally efficient time-adaptive aggregation (CETAA) module to model the time-varying characteristics of emotions. Experiments on RAVDESS, BAUM-1 s and eNTERFACE05 datasets also demonstrate that the proposed AataNet achieves better results.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

求助全文

约1分钟内获得全文去求助

来源期刊

IEEE Signal Processing Letters 工程技术-工程：电子与电气

CiteScore

7.40

自引率

12.80%

发文量

339

审稿时长

2.8 months

期刊介绍： The IEEE Signal Processing Letters is a monthly, archival publication designed to provide rapid dissemination of original, cutting-edge ideas and timely, significant contributions in signal, image, speech, language and audio processing. Papers published in the Letters can be presented within one year of their appearance in signal processing conferences such as ICASSP, GlobalSIP and ICIP, and also in several workshop organized by the Signal Processing Society.