网格工作流:一个灵活的网格故障处理框架

High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on Pub Date : 2003-06-22 DOI:10.1109/HPDC.2003.1210023

Soonwook Hwang, C. Kesselman

{"title":"网格工作流:一个灵活的网格故障处理框架","authors":"Soonwook Hwang, C. Kesselman","doi":"10.1109/HPDC.2003.1210023","DOIUrl":null,"url":null,"abstract":"The generic, heterogeneous, and dynamic nature of the grid requires a new from of failure recovery mechanism to address its unique requirements such as support for diverse failure handling strategies, separation of failure handling strategies from application codes, and user-defined exception handling. We here propose a grid workflow system (grid-WFS), a flexible failure handling framework for the grid, which addresses these grid-unique failure recovery requirements. Central to the framework is flexibility by the use of workflow structure as a high-level recovery policy specification. We show how this use of high-level workflow structure allows users to achieve failure recovery in a variety of ways depending on the requirements and constraints of their applications. We also demonstrate that this use of workflow structure enables users to not only rapidly prototype and investigate failure handling strategies, but also easily change them by simply modifying the encompassing workflow structure, while the application code remains intact. Finally, we present an experimental evaluation of our framework using a simulation, demonstrating the value of supporting multiple failure recovery techniques in grid systems to achieve high performance in the presence of failures.","PeriodicalId":430378,"journal":{"name":"High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2003-06-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"185","resultStr":"{\"title\":\"Grid workflow: a flexible failure handling framework for the grid\",\"authors\":\"Soonwook Hwang, C. Kesselman\",\"doi\":\"10.1109/HPDC.2003.1210023\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The generic, heterogeneous, and dynamic nature of the grid requires a new from of failure recovery mechanism to address its unique requirements such as support for diverse failure handling strategies, separation of failure handling strategies from application codes, and user-defined exception handling. We here propose a grid workflow system (grid-WFS), a flexible failure handling framework for the grid, which addresses these grid-unique failure recovery requirements. Central to the framework is flexibility by the use of workflow structure as a high-level recovery policy specification. We show how this use of high-level workflow structure allows users to achieve failure recovery in a variety of ways depending on the requirements and constraints of their applications. We also demonstrate that this use of workflow structure enables users to not only rapidly prototype and investigate failure handling strategies, but also easily change them by simply modifying the encompassing workflow structure, while the application code remains intact. Finally, we present an experimental evaluation of our framework using a simulation, demonstrating the value of supporting multiple failure recovery techniques in grid systems to achieve high performance in the presence of failures.\",\"PeriodicalId\":430378,\"journal\":{\"name\":\"High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2003-06-22\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"185\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/HPDC.2003.1210023\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/HPDC.2003.1210023","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 185

摘要

网格的通用性、异构性和动态性需要一种新的故障恢复机制来满足其独特的需求，例如支持多种故障处理策略、将故障处理策略与应用程序代码分离，以及用户定义的异常处理。本文提出一种网格工作流系统(grid- wfs)，它是一种灵活的网格故障处理框架，可以解决网格特有的故障恢复需求。该框架的核心是通过使用工作流结构作为高级恢复策略规范来实现灵活性。我们将展示高级工作流结构的使用如何允许用户根据其应用程序的需求和约束以各种方式实现故障恢复。我们还证明了这种工作流结构的使用使用户不仅可以快速地建立原型并调查故障处理策略，而且可以通过简单地修改包含的工作流结构来轻松地更改它们，同时应用程序代码保持完整。最后，我们使用模拟对我们的框架进行了实验评估，展示了在网格系统中支持多种故障恢复技术以在故障存在时实现高性能的价值。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Grid workflow: a flexible failure handling framework for the grid

The generic, heterogeneous, and dynamic nature of the grid requires a new from of failure recovery mechanism to address its unique requirements such as support for diverse failure handling strategies, separation of failure handling strategies from application codes, and user-defined exception handling. We here propose a grid workflow system (grid-WFS), a flexible failure handling framework for the grid, which addresses these grid-unique failure recovery requirements. Central to the framework is flexibility by the use of workflow structure as a high-level recovery policy specification. We show how this use of high-level workflow structure allows users to achieve failure recovery in a variety of ways depending on the requirements and constraints of their applications. We also demonstrate that this use of workflow structure enables users to not only rapidly prototype and investigate failure handling strategies, but also easily change them by simply modifying the encompassing workflow structure, while the application code remains intact. Finally, we present an experimental evaluation of our framework using a simulation, demonstrating the value of supporting multiple failure recovery techniques in grid systems to achieve high performance in the presence of failures.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

High Performance Distributed Computing, 2003. Proceedings. 12th IEEE International Symposium on

自引率

0.00%

发文量

期刊最新文献

Adaptive polling of grid resource monitors using a slacker coherence model Optimizing GridFTP through dynamic right-sizing Dynamic virtual clusters in a grid site manager Distributed pagerank for P2P systems Using views for customizing reusable components in component-based frameworks