Distinguishing Unseen from Seen for Generalized Zero-shot Learning

2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Pub Date : 2022-06-01 DOI:10.1109/CVPR52688.2022.00773

Hongzu Su, Jingjing Li, Zhi Chen, Lei Zhu, Ke Lu

{"title":"Distinguishing Unseen from Seen for Generalized Zero-shot Learning","authors":"Hongzu Su, Jingjing Li, Zhi Chen, Lei Zhu, Ke Lu","doi":"10.1109/CVPR52688.2022.00773","DOIUrl":null,"url":null,"abstract":"Generalized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Recognizing unseen classes as seen ones or vice versa often leads to poor performance in GZSL. Therefore, distinguishing seen and unseen domains is naturally an effective yet challenging solution for GZSL. In this paper, we present a novel method which leverages both visual and semantic modalities to distinguish seen and unseen categories. Specifically, our method deploys two variational autoencoders to generate latent representations for visual and semantic modalities in a shared latent space, in which we align latent representations of both modalities by Wasserstein distance and reconstruct two modalities with the representations of each other. In order to learn a clearer boundary between seen and unseen classes, we propose a two-stage training strategy which takes advantage of seen and unseen semantic descriptions and searches a threshold to separate seen and unseen visual samples. At last, a seen expert and an unseen expert are used for final classification. Extensive experiments on five widely used benchmarks verify that the proposed method can significantly improve the results of GZSL. For instance, our method correctly recognizes more than 99% samples when separating domains and improves the final classification accuracy from 72.6% to 82.9% on AWA1.","PeriodicalId":355552,"journal":{"name":"2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","volume":"756 ","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"10","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/CVPR52688.2022.00773","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 10

Abstract

Generalized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Recognizing unseen classes as seen ones or vice versa often leads to poor performance in GZSL. Therefore, distinguishing seen and unseen domains is naturally an effective yet challenging solution for GZSL. In this paper, we present a novel method which leverages both visual and semantic modalities to distinguish seen and unseen categories. Specifically, our method deploys two variational autoencoders to generate latent representations for visual and semantic modalities in a shared latent space, in which we align latent representations of both modalities by Wasserstein distance and reconstruct two modalities with the representations of each other. In order to learn a clearer boundary between seen and unseen classes, we propose a two-stage training strategy which takes advantage of seen and unseen semantic descriptions and searches a threshold to separate seen and unseen visual samples. At last, a seen expert and an unseen expert are used for final classification. Extensive experiments on five widely used benchmarks verify that the proposed method can significantly improve the results of GZSL. For instance, our method correctly recognizes more than 99% samples when separating domains and improves the final classification accuracy from 72.6% to 82.9% on AWA1.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

广义零射击学习中未见与已见的区分

广义零概率学习(GZSL)旨在识别在训练中可能没有看到类别的样本。将不可见类识别为可见类，反之亦然，通常会导致GZSL中的性能不佳。因此，区分可见域和不可见域自然是GZSL的一个有效但具有挑战性的解决方案。在本文中，我们提出了一种利用视觉和语义模式来区分可见和未见类别的新方法。具体来说，我们的方法部署了两个变分自编码器，在共享的潜在空间中生成视觉和语义模态的潜在表征，其中我们通过沃瑟斯坦距离对齐两种模态的潜在表征，并用彼此的表征重建两种模态。为了学习更清晰的可见类和不可见类之间的边界，我们提出了一种利用可见和不可见语义描述的两阶段训练策略，并搜索阈值来分离可见和不可见的视觉样本。最后，使用一个可见专家和一个不可见专家进行最终分类。在五个广泛使用的基准测试上进行的大量实验验证了该方法可以显著改善GZSL的结果。例如，我们的方法在分离域时正确识别了99%以上的样本，并将最终的分类准确率从AWA1上的72.6%提高到82.9%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

自引率

0.00%

发文量

期刊最新文献

Synthetic Aperture Imaging with Events and Frames PhotoScene: Photorealistic Material and Lighting Transfer for Indoor Scenes A Unified Model for Line Projections in Catadioptric Cameras with Rotationally Symmetric Mirrors Distinguishing Unseen from Seen for Generalized Zero-shot Learning Virtual Correspondence: Humans as a Cue for Extreme-View Geometry