Improving the Shuffle of Hadoop MapReduce

2013 IEEE 5th International Conference on Cloud Computing Technology and Science Pub Date : 2013-12-02 DOI:10.1109/CloudCom.2013.42

Jingui Li, Xuelian Lin, Xiaolong Cui, Yue Ye

引用次数: 18

Abstract

As an efficient parallel computing system based on MapReduce model, Hadoop is widely used for large-scale data analysis such as data mining, machine learning and scientific simulation. However, there are still some performance problems in MapReduce, especially the situation in the shuffle phase. In order to solve these problems, in this paper, a lightweight individual shuffle service component with more efficient I/O policy was proposed rather than the existing shuffle phase in MapReduce. We also describe how to implement the shuffle service in three steps: extract shuffle from reduce task as a shuffle task, reconstruct the shuffle task as a service and improve I/O scheduling policy on Map sides. Furthermore both simulated experiments and MapReduce job comparative studies are conducted to evaluate the performance of our improvements. The result reveals that our approach can decrease the whole job's execution time and make full use of cluster resources.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

改进Hadoop MapReduce的Shuffle

Hadoop作为一种基于MapReduce模型的高效并行计算系统，广泛应用于数据挖掘、机器学习、科学仿真等大规模数据分析。但是，在MapReduce中仍然存在一些性能问题，特别是shuffle阶段的情况。为了解决这些问题，本文提出了一种轻量级的单个shuffle服务组件，该组件具有更高效的I/O策略，而不是MapReduce中现有的shuffle阶段。我们还描述了如何分三步实现shuffle服务:从reduce任务中提取shuffle作为shuffle任务，重构shuffle任务作为服务，改进Map端的I/O调度策略。此外，还进行了模拟实验和MapReduce作业比较研究，以评估我们改进的性能。结果表明，该方法可以减少整个作业的执行时间，充分利用集群资源。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2013 IEEE 5th International Conference on Cloud Computing Technology and Science

自引率

0.00%

发文量

期刊最新文献

A Feasibility Study of Host-Level Contention Detection by Guest Virtual Machines Porting Grid Applications to the Cloud with Schlouder Towards Data Handling Requirements-Aware Cloud Computing Providing Desirable Data to Users When Integrating Wireless Sensor Networks with Mobile Cloud MELA: Monitoring and Analyzing Elasticity of Cloud Services