Handling data skew in parallel hash join computation using two-phase scheduling

Proceedings 1st International Conference on Algorithms and Architectures for Parallel Processing Pub Date : 1995-04-19 DOI:10.1109/ICAPP.1995.472237

Xiaofang Zhou, M. Orlowska

引用次数: 18

Abstract

A large number of parallel join algorithms has been proposed to maintain load-balancing in the presence of data skew. However, one important type of data skew-join product skew (JPS)-has been little studied. In this paper, a dynamic parallel join algorithm, which employs a two-phase scheduling procedure, is designed to handle the JPS problem. Two sets of scheduling heuristics are studied against various parameters. It is shown that many of the existing algorithms can be regarded as a special case of our algorithm, whose cost is based on the nature of data skew. While it can cope with JPS which other algorithms cannot approach, it can be as efficient as most existing algorithms when JPS does not exist.<>

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

使用两阶段调度处理并行哈希连接计算中的数据倾斜

为了在存在数据倾斜的情况下保持负载平衡，已经提出了大量的并行连接算法。然而，一种重要的数据倾斜类型-连接产品倾斜(JPS)-很少被研究。本文设计了一种采用两阶段调度过程的动态并行连接算法来处理JPS问题。针对不同的调度参数，研究了两组调度启发式算法。结果表明，现有的许多算法都可以看作是我们算法的一个特例，其代价取决于数据倾斜的性质。虽然它可以处理其他算法无法处理的JPS，但当JPS不存在时，它可以像大多数现有算法一样高效。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings 1st International Conference on Algorithms and Architectures for Parallel Processing

自引率

0.00%

发文量

期刊最新文献

Approximation algorithms for time constrained scheduling A general definition of deadlocks for distributed systems Vectoring the N-body problem on the CM-5 On deflection worm routing on meshes Variable tracking technique: a single-pass method to determine data dependence