Parallel Implementation of Finite-element Mesh Integration Algorithm on Many Integrated Core
Kou Da-zh · 2015
A C++ 3-D finite-element mesh integration algorithm was implemented and profiled on a heterogeneous Intel CPU/MIC architecture.By virtually programing in the offload mode[1]with explicit copies,a sequence of key element-wise operations are fully parallelized utilizing massive concurrency of OpenMP threads on MIC devices.It is remarkably demonstrated that,in the sense of overall run-time efficiency,one fully employed 3115 A MIC card outweighs two 8-core Intel XeonTME5-2670 CPUs.However,possibly owing to cache contention among physical threads on individual MIC core,scalability is somehow below an ideal level.Current test unveils a good chance of transplanting a full finite-element analysis code onto a multi-CPU nodes/multi-MIC devices platform based on this single-process multithread building block presented here.