Processing LSTM in memory using hybrid network expansion model

2017 IEEE International Workshop on Signal Processing Systems (SiPS) Pub Date : 2017-10-01 DOI:10.1109/SiPS.2017.8110011

Yu Gong, Tingting Xu, Bo Liu, Wei-qi Ge, Jinjiang Yang, Jun Yang, Longxing Shi

引用次数: 1

Abstract

With the rapidly increasing applications of deep learning, LSTM-RNNs are widely used. Meanwhile, the complex data dependence and intensive computation limit the performance of the accelerators. In this paper, we first proposed a hybrid network expansion model to exploit the finegrained data parallelism. Based on the model, we implemented a Reconfigurable Processing Unit(RPU) using Processing In Memory(PIM) units. Our work shows that the gates and cells in LSTM can be partitioned to fundamental operations and then recombined and mapped into heterogeneous computing components. The experimental results show that, implemented on 45nm CMOS process, the proposed RPU with size of 1.51 mm2 and power of 413 mw achieves 309 GOPS/W in power efficiency, and is 1.7 χ better than state-of-the-art reconfigurable architecture.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

使用混合网络扩展模型处理内存中的LSTM

随着深度学习应用的迅速增加，lstm - rnn得到了广泛的应用。同时，复杂的数据依赖性和密集的计算量限制了加速器的性能。本文首先提出了一种利用细粒度数据并行性的混合网络扩展模型。基于该模型，我们使用内存处理(PIM)单元实现了可重构处理单元(RPU)。我们的工作表明，LSTM中的门和单元可以划分为基本操作，然后重新组合并映射为异构计算组件。实验结果表明，在45nm CMOS工艺上实现的RPU尺寸为1.51 mm2，功耗为413 mw，功率效率为309 GOPS/W，比目前最先进的可重构架构提高1.7 χ。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2017 IEEE International Workshop on Signal Processing Systems (SiPS)

自引率

0.00%

发文量

期刊最新文献

Analysing the performance of divide-and-conquer sequential matrix diagonalisation for large broadband sensor arrays Design space exploration of dataflow-based Smith-Waterman FPGA implementations Hardware error correction using local syndromes A stochastic number representation for fully homomorphic cryptography Statistical analysis of Post-HEVC encoded videos