Medical Image Processing on Intel Parallel Frameworks
Jonathan Low · 2013
Increases in processor performance are now expected to mainly come from an increasing number of processing cores after clockspeeds reached a limit in 2005. Accompanying this are demands placed on software to become highly parallel and able to scale well to higher thread counts. Investments in time and effort on parallelising software have resulted in applications that exhibit excellent performance and scalability on multisocket CPU nodes. The Intel Xeon Phi product family aims to offer higher performance to those existing highly parallel applications by increasing the thread count well into the hundreds whilst allowing to keep the same programming model. This project follows the proposition of developing a parallel application that scales well on multicore CPU nodes with a view to achieving higher performance on an Intel Xeon Phi co-processor. This project aims to parallelise a medical imaging application from the Western General Hospital in Edinburgh used to help in the prediction of radiation-induced fibrosis for lung cancer patients after radiotherapy. We use the Intel Parallel Studio XE 2013 software development kit, utilising OpenMP for multithreading and the Intel Math Kernel Library to accelerate application performance on shared memory nodes. The project is then extended to examine the application performance on many more threads on an Intel Xeon Phi. By a change in the image filtering algorithm utilising Fast Fourier Transforms in addition to parallelisation, we obtain runtimes that are upto 168 times faster over the original code. This was achieved on dual socket Intel Xeon nodes with 48 and 64GB of memory. Running the same computational code on an Intel Xeon Phi resulted in a levelling of performance after more than sixty threads for the largest image we could run on the coprocessor. The 8GB memory limitation prevented the running of a number of threads required for high performance and the runtimes obtained on the Xeon Phi are higher than that achieved on the Intel Xeon compute-nodes.