STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition

IF 11.3 2区 计算机科学 Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE IEEE Transactions on Affective Computing Pub Date : 2024-10-07 DOI:10.1109/TAFFC.2024.3475729
Yi Chang;Zhao Ren;Zixing Zhang;Xin Jing;Kun Qian;Xi Shao;Bin Hu;Tanja Schultz;Björn W. Schuller
{"title":"STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition","authors":"Yi Chang;Zhao Ren;Zixing Zhang;Xin Jing;Kun Qian;Xi Shao;Bin Hu;Tanja Schultz;Björn W. Schuller","doi":"10.1109/TAFFC.2024.3475729","DOIUrl":null,"url":null,"abstract":"Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.","PeriodicalId":13131,"journal":{"name":"IEEE Transactions on Affective Computing","volume":"16 2","pages":"861-874"},"PeriodicalIF":11.3000,"publicationDate":"2024-10-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Affective Computing","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10706816/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0

Abstract

Speech contains rich information on the emotions of humans, and Speech Emotion Recognition (SER) has been an important topic in the area of human-computer interaction. The robustness of SER models is crucial, particularly in privacy-sensitive and reliability-demanding domains like private healthcare. Recently, the vulnerability of deep neural networks in the audio domain to adversarial attacks has become a popular area of research. However, prior works on adversarial attacks in the audio domain primarily rely on iterative gradient-based techniques, which are time-consuming and prone to overfitting the specific threat model. Furthermore, the exploration of sparse perturbations, which have the potential for better stealthiness, remains limited in the audio domain. To address these challenges, we propose a generator-based attack method to generate sparse and transferable adversarial examples to deceive SER models in an end-to-end and efficient manner. We evaluate our method on two widely-used SER datasets, Database of Elicited Mood in Speech (DEMoS) and Interactive Emotional dyadic MOtion CAPture (IEMOCAP), and demonstrate its ability to generate successful sparse adversarial examples in an efficient manner. Moreover, our generated adversarial examples exhibit model-agnostic transferability, enabling effective adversarial attacks on advanced victim models.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
STAA-Net:用于语音情感识别的稀疏可转移对抗攻击
语音包含了人类情感的丰富信息,语音情感识别(SER)已成为人机交互领域的一个重要课题。SER模型的健壮性至关重要,特别是在隐私敏感和可靠性要求高的领域,如私人医疗保健领域。近年来,音频领域的深度神经网络在对抗性攻击下的脆弱性已成为研究的热点。然而,先前针对音频领域的对抗性攻击的研究主要依赖于基于迭代梯度的技术,这种技术耗时且容易过度拟合特定的威胁模型。此外,对稀疏扰动的探索,有更好的隐身性的潜力,仍然局限于音频领域。为了解决这些挑战,我们提出了一种基于生成器的攻击方法,以端到端有效的方式生成稀疏和可转移的对抗示例来欺骗SER模型。我们在两个广泛使用的SER数据集(Database of Elicited Mood in Speech, DEMoS)和交互式情绪二元动作捕捉(Interactive Emotional dyadic MOtion CAPture, IEMOCAP)上评估了我们的方法,并证明了它能够高效地生成成功的稀疏对抗示例。此外,我们生成的对抗性示例显示出与模型无关的可转移性,从而能够对高级受害者模型进行有效的对抗性攻击。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
IEEE Transactions on Affective Computing
IEEE Transactions on Affective Computing COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE-COMPUTER SCIENCE, CYBERNETICS
CiteScore
15.00
自引率
6.20%
发文量
174
期刊介绍: The IEEE Transactions on Affective Computing is an international and interdisciplinary journal. Its primary goal is to share research findings on the development of systems capable of recognizing, interpreting, and simulating human emotions and related affective phenomena. The journal publishes original research on the underlying principles and theories that explain how and why affective factors shape human-technology interactions. It also focuses on how techniques for sensing and simulating affect can enhance our understanding of human emotions and processes. Additionally, the journal explores the design, implementation, and evaluation of systems that prioritize the consideration of affect in their usability. We also welcome surveys of existing work that provide new perspectives on the historical and future directions of this field.
期刊最新文献
STMAE-Few: A Spatial-Temporal Masked Autoencoder for Few-channel EEG-based Emotion Recognition AU-Guided Neural Prototype Trees for Interpretable Micro-Expression Recognition Bio-Behaviorally Inspired Artificial Mimosa for Ambient Haptic Sensing and Interaction Subject-Invariant Multimodal Affective Representation Learning Via Masked Cross-Modal Alignment Ternary Paradigm-Oriented and Hybrid-Graph Capsule Aggregation for Aspect-Based Sentiment Analysis
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1