CORC  > 软件研究所  > 软件所图书馆  > 会议论文
Enabling and scaling a global shallow-water atmospheric model on Tianhe-2
Xue, Wei (1) ; Yang, Chao (2) ; Fu, Haohuan (3) ; Wang, Xinliang (1) ; Xu, Yangtong (1) ; Gan, Lin (1) ; Lu, Yutong (5) ; Zhu, Xiaoqian (5)
2014
会议名称28th IEEE International Parallel and Distributed Processing Symposium, IPDPS 2014
会议日期May 19, 2014 - May 23, 2014
会议地点Phoenix, AZ, United states
页码745-754
中文摘要This paper presents a hybrid algorithm for the petascale global simulation of atmospheric dynamics on Tianhe-2, the world's current top-ranked supercomputer developed by China's National University of Defense Technology (NUDT). Tianhe-2 is equipped with both Intel Xeon CPUs and Intel Xeon Phi accelerators. A key idea of the hybrid algorithm is to enable flexible domain partition between an arbitrary number of processors and accelerators, so as to achieve a balanced and efficient utilization of the entire system. We also present an asynchronous and concurrent data transfer scheme to reduce the communication overhead between CPU and accelerators. The acceleration of our global atmospheric model is conducted to improve the use of the Intel MIC architecture. For the single-node test on Tianhe-2 against two Intel Ivy Bridge CPUs (24 cores), we can achieve 2.07x, 3.18x, and 4.35x speedups when using one, two, and three Intel Xeon Phi accelerators respectively. The average performance gain from SIMD vectorization on the Intel Xeon Phi processors is around 5x (out of the 8x theoretical case). Based on successful computation-communication overlapping, large-scale tests indicate that a nearly ideal weak-scaling efficiency of 93.5% is obtained when we gradually increase the number of nodes from 6 to 8,664 (nearly 1.7 million cores). In the strong-scaling test, the parallel efficiency is about 77% when the number of nodes increases from 1,536 to 8,664 for a fixed 65,664 × 5,664 × 6 mesh with 77.6 billion unknowns. © 2014 IEEE.
英文摘要This paper presents a hybrid algorithm for the petascale global simulation of atmospheric dynamics on Tianhe-2, the world's current top-ranked supercomputer developed by China's National University of Defense Technology (NUDT). Tianhe-2 is equipped with both Intel Xeon CPUs and Intel Xeon Phi accelerators. A key idea of the hybrid algorithm is to enable flexible domain partition between an arbitrary number of processors and accelerators, so as to achieve a balanced and efficient utilization of the entire system. We also present an asynchronous and concurrent data transfer scheme to reduce the communication overhead between CPU and accelerators. The acceleration of our global atmospheric model is conducted to improve the use of the Intel MIC architecture. For the single-node test on Tianhe-2 against two Intel Ivy Bridge CPUs (24 cores), we can achieve 2.07x, 3.18x, and 4.35x speedups when using one, two, and three Intel Xeon Phi accelerators respectively. The average performance gain from SIMD vectorization on the Intel Xeon Phi processors is around 5x (out of the 8x theoretical case). Based on successful computation-communication overlapping, large-scale tests indicate that a nearly ideal weak-scaling efficiency of 93.5% is obtained when we gradually increase the number of nodes from 6 to 8,664 (nearly 1.7 million cores). In the strong-scaling test, the parallel efficiency is about 77% when the number of nodes increases from 1,536 to 8,664 for a fixed 65,664 × 5,664 × 6 mesh with 77.6 billion unknowns. © 2014 IEEE.
收录类别EI
会议录出版地IEEE Computer Society
语种英语
ISSN号15302075
ISBN号9780769552071
内容类型会议论文
源URL[http://ir.iscas.ac.cn/handle/311060/16636]  
专题软件研究所_软件所图书馆_会议论文
推荐引用方式
GB/T 7714
Xue, Wei ,Yang, Chao ,Fu, Haohuan ,et al. Enabling and scaling a global shallow-water atmospheric model on Tianhe-2[C]. 见:28th IEEE International Parallel and Distributed Processing Symposium, IPDPS 2014. Phoenix, AZ, United states. May 19, 2014 - May 23, 2014.
个性服务
查看访问统计
相关权益政策
暂无数据
收藏/分享
所有评论 (0)
暂无评论
 

除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。


©版权所有 ©2017 CSpace - Powered by CSpace