高可用性HPC系统服务的主/主复制

First International Conference on Availability, Reliability and Security (ARES'06) Pub Date : 2006-04-20 DOI:10.1109/ARES.2006.23

C. Engelmann, S. Scott, C. Leangsuksun, Xubin He

{"title":"高可用性HPC系统服务的主/主复制","authors":"C. Engelmann, S. Scott, C. Leangsuksun, Xubin He","doi":"10.1109/ARES.2006.23","DOIUrl":null,"url":null,"abstract":"Today's high performance computing systems have several reliability deficiencies resulting in availability and serviceability issues. Head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. This paper introduces two distinct replication methods (internal and external) for providing symmetric active/active high availability for multiple head and service nodes running in virtual synchrony. It presents a comparison of both methods in terms of expected correctness, ease-of-use and performance based on early results from ongoing work in providing symmetric active/active high availability for two HPC system services (TORQUE and PVFS metadata server). It continues with a short description of a distributed mutual exclusion algorithm and a brief statement regarding the handling of Byzantine failures. This paper concludes with an overview of past and ongoing work, and a short summary of the presented research.","PeriodicalId":106780,"journal":{"name":"First International Conference on Availability, Reliability and Security (ARES'06)","volume":"49 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2006-04-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"14","resultStr":"{\"title\":\"Active/active replication for highly available HPC system services\",\"authors\":\"C. Engelmann, S. Scott, C. Leangsuksun, Xubin He\",\"doi\":\"10.1109/ARES.2006.23\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Today's high performance computing systems have several reliability deficiencies resulting in availability and serviceability issues. Head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. This paper introduces two distinct replication methods (internal and external) for providing symmetric active/active high availability for multiple head and service nodes running in virtual synchrony. It presents a comparison of both methods in terms of expected correctness, ease-of-use and performance based on early results from ongoing work in providing symmetric active/active high availability for two HPC system services (TORQUE and PVFS metadata server). It continues with a short description of a distributed mutual exclusion algorithm and a brief statement regarding the handling of Byzantine failures. This paper concludes with an overview of past and ongoing work, and a short summary of the presented research.\",\"PeriodicalId\":106780,\"journal\":{\"name\":\"First International Conference on Availability, Reliability and Security (ARES'06)\",\"volume\":\"49 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2006-04-20\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"14\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"First International Conference on Availability, Reliability and Security (ARES'06)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ARES.2006.23\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"First International Conference on Availability, Reliability and Security (ARES'06)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ARES.2006.23","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 14

摘要

当今的高性能计算系统存在一些可靠性缺陷，从而导致可用性和可维护性问题。头节点和服务节点代表了整个系统的单点故障和控制，因为它们使系统在故障发生时无法访问和管理，直到修复，从而导致大量停机时间。本文介绍了两种不同的复制方法(内部和外部)，用于为以虚拟同步方式运行的多个头部和服务节点提供对称的主动/主动高可用性。本文根据正在进行的为两个HPC系统服务(TORQUE和PVFS元数据服务器)提供对称主动/主动高可用性的工作的早期结果，对两种方法在预期正确性、易用性和性能方面进行了比较。接着简要介绍了分布式互斥算法，并简要介绍了拜占庭故障的处理。本文总结了过去和正在进行的工作，并对所提出的研究进行了简短的总结。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Active/active replication for highly available HPC system services

Today's high performance computing systems have several reliability deficiencies resulting in availability and serviceability issues. Head and service nodes represent a single point of failure and control for an entire system as they render it inaccessible and unmanageable in case of a failure until repair, causing a significant downtime. This paper introduces two distinct replication methods (internal and external) for providing symmetric active/active high availability for multiple head and service nodes running in virtual synchrony. It presents a comparison of both methods in terms of expected correctness, ease-of-use and performance based on early results from ongoing work in providing symmetric active/active high availability for two HPC system services (TORQUE and PVFS metadata server). It continues with a short description of a distributed mutual exclusion algorithm and a brief statement regarding the handling of Byzantine failures. This paper concludes with an overview of past and ongoing work, and a short summary of the presented research.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

First International Conference on Availability, Reliability and Security (ARES'06)

自引率

0.00%

发文量

期刊最新文献

Inter-domains security management (IDSM) model for IP multimedia subsystem (IMS) Securing DNS services through system self cleansing and hardware enhancements No risk is unsafe: simulated results on dependability of complementary currencies Quality of password management policy Recovery mechanism of cooperative process chain in grid