Progressive Masking Oriented Self-Taught Learning for Occluded Facial Expression Recognition

IF 9.8 2区计算机科学 Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE IEEE Transactions on Affective Computing Pub Date : 2025-02-25 DOI:10.1109/TAFFC.2025.3544677

Bin Kang;Shuangshuang Wang;Zongyu Wang;Xin Li;Haie Dou;Lei Wang;Zhijie Xia

{"title":"Progressive Masking Oriented Self-Taught Learning for Occluded Facial Expression Recognition","authors":"Bin Kang;Shuangshuang Wang;Zongyu Wang;Xin Li;Haie Dou;Lei Wang;Zhijie Xia","doi":"10.1109/TAFFC.2025.3544677","DOIUrl":null,"url":null,"abstract":"Self-taught learning (STL) is a promising solution that reduces the performance gap between weakly supervised and fully supervised learning for easily accessible, label-free images. The success of traditional STL solutions relies on the assumption that the target appearance is completely visible and well-defined. In real-world facial expression recognition scenarios, however, saliency regions are often partially occluded, which significantly hampers the generalization capability of STL methods. Nevertheless, few studies have investigated the impact of occlusion on STL. In this paper, we propose an interweaved autoencoder network for weakly supervised facial expression recognition in occlusion scenarios. The key innovation of our network lies in the Residual Connection Union (RCU) blocks that can integrate the Convolutional Neural Network (CNN) and Transformer layers into a multi-scale structure. The RCU enables a progressive masking strategy to accurately identify and focus on contributive yet often overlooked image patches by analyzing the relationships among region-level target representations. In addition, we introduce a self-knowledge distillation module for the effective training of the proposed autoencoder network. Extensive experiments are conducted on four public datasets to demonstrate the superiority of our method over related works.","PeriodicalId":13131,"journal":{"name":"IEEE Transactions on Affective Computing","volume":"16 3","pages":"1277-1289"},"PeriodicalIF":9.8000,"publicationDate":"2025-02-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Affective Computing","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10902013/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}

引用次数: 0

Abstract

Self-taught learning (STL) is a promising solution that reduces the performance gap between weakly supervised and fully supervised learning for easily accessible, label-free images. The success of traditional STL solutions relies on the assumption that the target appearance is completely visible and well-defined. In real-world facial expression recognition scenarios, however, saliency regions are often partially occluded, which significantly hampers the generalization capability of STL methods. Nevertheless, few studies have investigated the impact of occlusion on STL. In this paper, we propose an interweaved autoencoder network for weakly supervised facial expression recognition in occlusion scenarios. The key innovation of our network lies in the Residual Connection Union (RCU) blocks that can integrate the Convolutional Neural Network (CNN) and Transformer layers into a multi-scale structure. The RCU enables a progressive masking strategy to accurately identify and focus on contributive yet often overlooked image patches by analyzing the relationships among region-level target representations. In addition, we introduce a self-knowledge distillation module for the effective training of the proposed autoencoder network. Extensive experiments are conducted on four public datasets to demonstrate the superiority of our method over related works.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

面向渐进式掩蔽的封闭面部表情识别自学

自学（STL）是一种很有前途的解决方案，它可以减少弱监督学习和完全监督学习之间的性能差距，用于易于访问的无标签图像。传统STL解决方案的成功依赖于目标外观完全可见且定义良好的假设。然而，在真实的面部表情识别场景中，显著区域往往被部分遮挡，这严重影响了STL方法的泛化能力。然而，很少有研究探讨闭塞对STL的影响。在本文中，我们提出了一种用于弱监督遮挡场景下面部表情识别的交织自编码器网络。我们的网络的关键创新在于残余连接联盟（RCU）块，它可以将卷积神经网络（CNN）和变压器层集成到一个多尺度结构中。RCU支持渐进式掩蔽策略，通过分析区域级目标表示之间的关系，准确地识别和关注有贡献但经常被忽视的图像补丁。此外，我们还引入了一个自知识蒸馏模块来有效地训练所提出的自编码器网络。在四个公共数据集上进行了大量的实验，以证明我们的方法优于相关工作。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

IEEE Transactions on Affective Computing COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE-COMPUTER SCIENCE, CYBERNETICS

CiteScore

15.00

自引率

6.20%

发文量

174

期刊介绍： The IEEE Transactions on Affective Computing is an international and interdisciplinary journal. Its primary goal is to share research findings on the development of systems capable of recognizing, interpreting, and simulating human emotions and related affective phenomena. The journal publishes original research on the underlying principles and theories that explain how and why affective factors shape human-technology interactions. It also focuses on how techniques for sensing and simulating affect can enhance our understanding of human emotions and processes. Additionally, the journal explores the design, implementation, and evaluation of systems that prioritize the consideration of affect in their usability. We also welcome surveys of existing work that provide new perspectives on the historical and future directions of this field.

期刊最新文献

CAST-Phys: Contactless Affective States Through Physiological Signals Database The MSP-Podcast Corpus Graph-Based Representation Learning with Beta Uncertainty for Enhanced Multimodal Emotion Recognition SpotFormer: Multi-Scale Spatio-Temporal Transformer for Facial Expression Spotting Weakly Supervised Learning for Facial Affective Behavior Analysis: a Review