Characterizing Loop-Level Communication Patterns in Shared Memory

2015 44th International Conference on Parallel Processing Pub Date : 2015-09-01 DOI:10.1109/ICPP.2015.85

Arya Mazaheri, A. Jannesari, Abdolreza Mirzaei, F. Wolf

{"title":"Characterizing Loop-Level Communication Patterns in Shared Memory","authors":"Arya Mazaheri, A. Jannesari, Abdolreza Mirzaei, F. Wolf","doi":"10.1109/ICPP.2015.85","DOIUrl":null,"url":null,"abstract":"Communication patterns extracted from parallel programs can provide a valuable source of information for parallel pattern detection, application auto-tuning, and runtime workload scheduling on heterogeneous systems. Once identified, such patterns can help find the most promising optimizations. Communication patterns can be detected using different methods, including sandbox simulation, memory profiling, and hardware counter analysis. However, these analyses usually suffer from high runtime and memory overhead, necessitating a trade off between accuracy and resource consumption. More importantly, none of the existing methods exploit fine-grained communication patterns on the level of individual code regions. In this paper, we present an efficient tool based on Disco PoP profiler that characterizes the communication pattern of every hotspot in a shared-memory application. With the aid of static and dynamic code analysis, it produces a nested structure of communication patterns based on program's loops. By employing asymmetric signature memory, the runtime overhead is around 225× while the required amount of memory remains fixed. In comparison with other profilers, the proposed method is efficient enough to be used with real world applications.","PeriodicalId":423007,"journal":{"name":"2015 44th International Conference on Parallel Processing","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2015-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2015 44th International Conference on Parallel Processing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICPP.2015.85","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

Abstract

Communication patterns extracted from parallel programs can provide a valuable source of information for parallel pattern detection, application auto-tuning, and runtime workload scheduling on heterogeneous systems. Once identified, such patterns can help find the most promising optimizations. Communication patterns can be detected using different methods, including sandbox simulation, memory profiling, and hardware counter analysis. However, these analyses usually suffer from high runtime and memory overhead, necessitating a trade off between accuracy and resource consumption. More importantly, none of the existing methods exploit fine-grained communication patterns on the level of individual code regions. In this paper, we present an efficient tool based on Disco PoP profiler that characterizes the communication pattern of every hotspot in a shared-memory application. With the aid of static and dynamic code analysis, it produces a nested structure of communication patterns based on program's loops. By employing asymmetric signature memory, the runtime overhead is around 225× while the required amount of memory remains fixed. In comparison with other profilers, the proposed method is efficient enough to be used with real world applications.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

共享内存中环路级通信模式的表征

从并行程序中提取的通信模式可以为异构系统上的并行模式检测、应用程序自动调优和运行时工作负载调度提供有价值的信息源。一旦确定，这样的模式可以帮助找到最有前途的优化。可以使用不同的方法检测通信模式，包括沙箱模拟、内存分析和硬件计数器分析。然而，这些分析通常受到高运行时和内存开销的影响，需要在准确性和资源消耗之间进行权衡。更重要的是，现有的方法都没有利用单个代码区域级别上的细粒度通信模式。在本文中，我们提出了一个基于Disco PoP分析器的高效工具来表征共享内存应用程序中每个热点的通信模式。借助静态和动态代码分析，生成基于程序循环的通信模式嵌套结构。通过使用非对称签名内存，运行时开销约为225x，而所需的内存量保持不变。与其他分析器相比，所提出的方法是有效的，足以在实际应用中使用。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2015 44th International Conference on Parallel Processing

自引率

0.00%

发文量