首页 | 本学科首页   官方微博 | 高级检索  
     

面向分布式图计算作业的容错技术研究综述
引用本文:张程博,李影,贾统.面向分布式图计算作业的容错技术研究综述[J].软件学报,2021,32(7):2078-2102.
作者姓名:张程博  李影  贾统
作者单位:北京大学 软件与微电子学院, 北京 海淀 102600;北京大学 软件与微电子学院, 北京 海淀 102600;北京大学 软件工程国家工程研究中心, 北京 100871;北京大学 信息科学技术学院, 北京 100871
基金项目:广东省重点领域研发计划(NO.2020B010164003)
摘    要:随着图数据规模的日益庞大和图计算作业的日益复杂,图计算的分布化成为必然趋势.然而图计算作业在运行过程中面临着分布式图计算系统内外各种来源的非确定性所带来的严峻的可靠性问题.本文首先分析了分布式图计算框架中不确定性因素和不同类型图计算作业的鲁棒性,并提出了基于成本、效率和质量三个维度的面向分布式图计算作业的容错技术评估框架,然后分别对分布式图计算的四种容错机制——基于检查点的容错、基于日志的容错、基于复制的容错、基于算法补偿的容错等机制结合国内外相关工作做了深入地分析、评估和比较.最后对未来的研究方向做了展望.

关 键 词:图数据  故障和失效  分布式图计算  容错机制  非确定性软件系统
收稿时间:2020/9/15 0:00:00
修稿时间:2020/10/26 0:00:00

Survey of State-of-the-art Fault Tolerance for Distributed Graph Processing Jobs
ZHANG Cheng-Bo,LI Ying,JIA Tong.Survey of State-of-the-art Fault Tolerance for Distributed Graph Processing Jobs[J].Journal of Software,2021,32(7):2078-2102.
Authors:ZHANG Cheng-Bo  LI Ying  JIA Tong
Affiliation:School of Software and Microelectronics, Peking University, Beijing 102600, China;School of Software and Microelectronics, Peking University, Beijing 102600, China;National Engineering Research Center for Software Engineering, Peking University, Beijing 100871, China; School of Electronics and Computer Science, Peking University, Beijing 100871, China
Abstract:As the growth of graph data scale and complexity of graph processing, the trend of distributed graph processing shall be inevitable. However, graph processing jobs run with severe reliability problems caused by the uncertainty originated from inside and outside the distributed graph processing system. This study first analyzes the uncertainty factors of the distributed graph processing frameworks and the robustness of different types of graph processing jobs; then proposes an evaluation framework of fault tolerance for distributed graph processing based on cost, efficiency, and quality of fault tolerance. This study also analyzes, evaluates, and compares the four fault-tolerant mechanisms of distributed graph processing-checkpointing based fault tolerance, logging based fault tolerance, replication based fault tolerance, and algorithm compensation based fault tolerance-combining related researches. Finally, the direction of future researches is prospected.
Keywords:graph data  fault and failure  distributed graph processing  fault tolerance  uncertainty software system
点击此处可从《软件学报》浏览原始摘要信息
点击此处可从《软件学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号