首页 | 本学科首页   官方微博 | 高级检索  
     


Highly Optimized Code Generation for Stencil Codes with Computation Reuse for GPUs
Authors:Wen-Jing Ma  Kan Gao  Guo-Ping Long
Affiliation:1.Laboratory of Parallel Software and Computing Science, Institute of Software, Chinese Academy of Sciences,Beijing,China;2.State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences,Beijing,China;3.Information Center, China Association for Science and Technology,Beijing,China
Abstract:Computation reuse is known as an effective optimization technique. However, due to the complexity of modern GPU architectures, there is yet not enough understanding regarding the intriguing implications of the interplay of computation reuse and hardware specifics on application performance. In this paper, we propose an automatic code generator for a class of stencil codes with inherent computation reuse on GPUs. For such applications, the proper reuse of intermediate results, combined with careful register and on-chip local memory usage, has profound implications on performance. Current state of the art does not address this problem in depth, partially due to the lack of a good program representation that can expose all potential computation reuse. In this paper, we leverage the computation overlap graph (COG), a simple representation of data dependence and data reuse with “element view”, to expose potential reuse opportunities. Using COG, we propose a portable code generation and tuning framework for GPUs. Compared with current state-of-the-art code generators, our experimental results show up to 56.7 % performance improvement on modern GPUs such as NVIDIA C2050.
Keywords:
本文献已被 SpringerLink 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号