Bi-SAN-CAP:图像标题的双向自注意

2019 Digital Image Computing: Techniques and Applications (DICTA) Pub Date : 2019-12-01 DOI:10.1109/DICTA47822.2019.8946003

Md. Zakir Hossain, F. Sohel, M. F. Shiratuddin, Hamid Laga, Bennamoun

{"title":"Bi-SAN-CAP:图像标题的双向自注意","authors":"Md. Zakir Hossain, F. Sohel, M. F. Shiratuddin, Hamid Laga, Bennamoun","doi":"10.1109/DICTA47822.2019.8946003","DOIUrl":null,"url":null,"abstract":"In a typical image captioning pipeline, a Convolutional Neural Network (CNN) is used as the image encoder and Long Short-Term Memory (LSTM) as the language decoder. LSTM with attention mechanism has shown remarkable performance on sequential data including image captioning. LSTM can retain long-range dependency of sequential data. However, it is hard to parallelize the computations of LSTM because of its inherent sequential characteristics. In order to address this issue, recent works have shown benefits in using self-attention, which is highly parallelizable without requiring any temporal dependencies. However, existing techniques apply attention only in one direction to compute the context of the words. We propose an attention mechanism called Bi-directional Self-Attention (Bi-SAN) for image captioning. It computes attention both in forward and backward directions. It achieves high performance comparable to state-of-the-art methods.","PeriodicalId":6696,"journal":{"name":"2019 Digital Image Computing: Techniques and Applications (DICTA)","volume":"12 1","pages":"1-7"},"PeriodicalIF":0.0000,"publicationDate":"2019-12-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"10","resultStr":"{\"title\":\"Bi-SAN-CAP: Bi-Directional Self-Attention for Image Captioning\",\"authors\":\"Md. Zakir Hossain, F. Sohel, M. F. Shiratuddin, Hamid Laga, Bennamoun\",\"doi\":\"10.1109/DICTA47822.2019.8946003\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In a typical image captioning pipeline, a Convolutional Neural Network (CNN) is used as the image encoder and Long Short-Term Memory (LSTM) as the language decoder. LSTM with attention mechanism has shown remarkable performance on sequential data including image captioning. LSTM can retain long-range dependency of sequential data. However, it is hard to parallelize the computations of LSTM because of its inherent sequential characteristics. In order to address this issue, recent works have shown benefits in using self-attention, which is highly parallelizable without requiring any temporal dependencies. However, existing techniques apply attention only in one direction to compute the context of the words. We propose an attention mechanism called Bi-directional Self-Attention (Bi-SAN) for image captioning. It computes attention both in forward and backward directions. It achieves high performance comparable to state-of-the-art methods.\",\"PeriodicalId\":6696,\"journal\":{\"name\":\"2019 Digital Image Computing: Techniques and Applications (DICTA)\",\"volume\":\"12 1\",\"pages\":\"1-7\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-12-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"10\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2019 Digital Image Computing: Techniques and Applications (DICTA)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/DICTA47822.2019.8946003\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 Digital Image Computing: Techniques and Applications (DICTA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/DICTA47822.2019.8946003","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 10

摘要

在典型的图像字幕管道中，使用卷积神经网络(CNN)作为图像编码器，使用长短期记忆(LSTM)作为语言解码器。具有注意机制的LSTM在包括图像字幕在内的序列数据上表现出了显著的性能。LSTM可以保留序列数据的长期依赖关系。然而，由于LSTM固有的序列特性，其计算难以并行化。为了解决这个问题，最近的研究显示了使用自我关注的好处，它是高度并行化的，不需要任何时间依赖性。然而，现有的技术只在一个方向上应用注意力来计算单词的上下文。我们提出了一种称为双向自注意(Bi-SAN)的图像字幕注意机制。它计算向前和向后方向的注意力。它实现了与最先进的方法相媲美的高性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Bi-SAN-CAP: Bi-Directional Self-Attention for Image Captioning

In a typical image captioning pipeline, a Convolutional Neural Network (CNN) is used as the image encoder and Long Short-Term Memory (LSTM) as the language decoder. LSTM with attention mechanism has shown remarkable performance on sequential data including image captioning. LSTM can retain long-range dependency of sequential data. However, it is hard to parallelize the computations of LSTM because of its inherent sequential characteristics. In order to address this issue, recent works have shown benefits in using self-attention, which is highly parallelizable without requiring any temporal dependencies. However, existing techniques apply attention only in one direction to compute the context of the words. We propose an attention mechanism called Bi-directional Self-Attention (Bi-SAN) for image captioning. It computes attention both in forward and backward directions. It achieves high performance comparable to state-of-the-art methods.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2019 Digital Image Computing: Techniques and Applications (DICTA)

自引率

0.00%

发文量