通过反向计算利用信息局部性实现高性能

2011 23rd International Symposium on Computer Architecture and High Performance Computing Pub Date : 2011-10-26 DOI:10.1109/SBAC-PAD.2011.10

Mouad Bahi, C. Eisenbeis

{"title":"通过反向计算利用信息局部性实现高性能","authors":"Mouad Bahi, C. Eisenbeis","doi":"10.1109/SBAC-PAD.2011.10","DOIUrl":null,"url":null,"abstract":"In this paper we present performance results for our register rematerialization technique based on reverse recomputing. Rematerialization adds instructions and we show on one specifically designed example that reverse computing alleviates the impact of these additional instructions on performance. We also show how thread parallelism may be optimized on GPUs by performing register allocation with reverse recomputing that increases the number of threads per Streaming Multiprocessor (SM). This is done on the main kernel of Lattice Quantum Chromo Dynamics (LQCD) simulation program where we gain a 10.84% speedup.","PeriodicalId":390734,"journal":{"name":"2011 23rd International Symposium on Computer Architecture and High Performance Computing","volume":"3 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2011-10-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"High Performance by Exploiting Information Locality through Reverse Computing\",\"authors\":\"Mouad Bahi, C. Eisenbeis\",\"doi\":\"10.1109/SBAC-PAD.2011.10\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper we present performance results for our register rematerialization technique based on reverse recomputing. Rematerialization adds instructions and we show on one specifically designed example that reverse computing alleviates the impact of these additional instructions on performance. We also show how thread parallelism may be optimized on GPUs by performing register allocation with reverse recomputing that increases the number of threads per Streaming Multiprocessor (SM). This is done on the main kernel of Lattice Quantum Chromo Dynamics (LQCD) simulation program where we gain a 10.84% speedup.\",\"PeriodicalId\":390734,\"journal\":{\"name\":\"2011 23rd International Symposium on Computer Architecture and High Performance Computing\",\"volume\":\"3 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2011-10-26\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2011 23rd International Symposium on Computer Architecture and High Performance Computing\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SBAC-PAD.2011.10\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2011 23rd International Symposium on Computer Architecture and High Performance Computing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SBAC-PAD.2011.10","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

本文给出了基于反向重计算的寄存器重物化技术的性能结果。重新物质化增加了指令，我们在一个专门设计的示例中展示了反向计算减轻了这些额外指令对性能的影响。我们还展示了如何通过使用反向重计算执行寄存器分配来优化gpu上的线程并行性，这增加了每个流式多处理器(SM)的线程数量。这是在Lattice Quantum Chromo Dynamics (LQCD)模拟程序的主内核上完成的，我们获得了10.84%的加速。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

High Performance by Exploiting Information Locality through Reverse Computing

In this paper we present performance results for our register rematerialization technique based on reverse recomputing. Rematerialization adds instructions and we show on one specifically designed example that reverse computing alleviates the impact of these additional instructions on performance. We also show how thread parallelism may be optimized on GPUs by performing register allocation with reverse recomputing that increases the number of threads per Streaming Multiprocessor (SM). This is done on the main kernel of Lattice Quantum Chromo Dynamics (LQCD) simulation program where we gain a 10.84% speedup.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2011 23rd International Symposium on Computer Architecture and High Performance Computing

自引率

0.00%

发文量

期刊最新文献

A Power-Efficient Co-designed Out-of-Order Processor A Metadata Cluster Based on OSD+ Devices The Experience in Designing and Building the High Performance Cluster Netuno Predictive and Distributed Routing Balancing on High-Speed Cluster Networks Data Parallelism for Belief Propagation in Factor Graphs