Fast crash recovery in RAMCloud

Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles Pub Date : 2011-10-23 DOI:10.1145/2043556.2043560

Diego Ongaro, Stephen M. Rumble, Ryan Stutsman, J. Ousterhout, M. Rosenblum

引用次数: 375

Abstract

RAMCloud is a DRAM-based storage system that provides inexpensive durability and availability by recovering quickly after crashes, rather than storing replicas in DRAM. RAMCloud scatters backup data across hundreds or thousands of disks, and it harnesses hundreds of servers in parallel to reconstruct lost data. The system uses a log-structured approach for all its data, in DRAM as well as on disk: this provides high performance both during normal operation and during recovery. RAMCloud employs randomized techniques to manage the system in a scalable and decentralized fashion. In a 60-node cluster, RAMCloud recovers 35 GB of data from a failed server in 1.6 seconds. Our measurements suggest that the approach will scale to recover larger memory sizes (64 GB or more) in less time with larger clusters.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

快速崩溃恢复在RAMCloud

RAMCloud是一种基于DRAM的存储系统，它通过在崩溃后快速恢复提供廉价的持久性和可用性，而不是将副本存储在DRAM中。RAMCloud将备份数据分散在数百或数千个磁盘上，并并行利用数百台服务器来重建丢失的数据。该系统对其所有数据使用日志结构方法，存储在DRAM和磁盘上:这在正常操作和恢复期间都提供了高性能。RAMCloud采用随机化技术以可扩展和分散的方式管理系统。在60个节点的集群中，RAMCloud在1.6秒内从故障服务器恢复35gb的数据。我们的测量表明，该方法可以在更大的集群中在更短的时间内恢复更大的内存大小(64 GB或更多)。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles

自引率

0.00%

发文量