首页 | 本学科首页   官方微博 | 高级检索  
     


An uncoordinated asynchronous checkpointing model for hierarchical scientific workflows
Authors:Rafael Tolosana-Calasanz  José Ángel Bañares  Pedro Álvarez  Joaquín Ezpeleta  Omer Rana
Affiliation:1. Computer Science and Systems Engineering Department, Aragón Institute of Engineering Research (I3A), Universidad de Zaragoza, Spain;2. School of Computer Science, Cardiff University, United Kingdom
Abstract:Scientific workflow systems often operate in unreliable environments, and have accordingly incorporated different fault tolerance techniques. One of them is the checkpointing technique combined with its corresponding rollback recovery process. Different checkpointing schemes have been developed and at various levels: task- (or activity-) level and workflow-level. At workflow-level, the usually adopted approach is to establish a checkpointing frequency in the system which determines the moment at which a global workflow checkpoint – a snapshot of the whole workflow enactment state at normal execution (without failures) – has to be accomplished. We describe an alternative workflow-level checkpointing scheme and its corresponding rollback recovery process for hierarchical scientific workflows in which every workflow node in the hierarchy accomplishes its own local checkpoint autonomously and in an uncoordinated way after its enactment. In contrast to other proposals, we utilise the Reference net formalism for expressing the scheme. Reference nets are a particular type of Petri nets which can more effectively provide the abstractions to support and to express hierarchical workflows and their dynamic adaptability.
Keywords:
本文献已被 ScienceDirect 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号