Visual Story Ordering with a Bidirectional Writer

Proceedings of the 2020 International Conference on Multimedia Retrieval Pub Date : 2020-06-08 DOI:10.1145/3372278.3390735

Wei-Rou Lin, Hen-Hsen Huang, Hsin-Hsi Chen

引用次数: 0

Abstract

This paper introduces visual story ordering, a challenging task in which images and text are ordered in a visual story jointly. We propose a neural network model based on the reader-processor-writer architecture with a self-attention mechanism. A novel bidirectional decoder is further proposed with bidirectional beam search. Experimental results show the effectiveness of the approach. The information gained from multimodal learning is presented and discussed. We also find that the proposed embedding narrows the distance between images and their corresponding story sentences, even though we do not align the two modalities explicitly. As it addresses a general issue in generative models, the proposed bidirectional inference mechanism applies to a variety of applications.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

使用双向书写器的视觉故事排序

视觉故事排序是一项具有挑战性的任务，其中图像和文本在视觉故事中共同排序。提出了一种具有自关注机制的基于读-处理器-写体系结构的神经网络模型。进一步提出了一种新型双向解码器，采用双向波束搜索。实验结果表明了该方法的有效性。介绍并讨论了从多模态学习中获得的信息。我们还发现，即使我们没有明确地对齐两种模式，所提出的嵌入也缩小了图像与其对应的故事句子之间的距离。由于它解决了生成模型中的一般问题，因此所提出的双向推理机制适用于各种应用。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the 2020 International Conference on Multimedia Retrieval

自引率

0.00%

发文量

期刊最新文献

Music Tower Blocks: Multi-Faceted Exploration Interface for Web-Scale Music Access Deep Semantic-Alignment Hashing for Unsupervised Cross-Modal Retrieval Urban Movie Map for Walkers: Route View Synthesis using 360° Videos ICDAR'20: Intelligent Cross-Data Analysis and Retrieval An Interactive Multimodal Retrieval System for Memory Assistant and Life Organized Support