CPU and GPU Parallelization of Sparsity Driven Despeckling in SAR Images
Caner Ozcan, Baha Şen, Fatih Nar · AYBU AVESIS · 2016
SAR remote sensing systems provide high-resolution images ofthe earth's surface as seen from airborne or satellite platforms. SAR imaginghas a growing interest in remote sensing applications in many fields with theability to generate high-resolution radar image regardless of weatherconditions and sunlight illumination. However, due to coherent processing ofreceived signals, SAR images are degraded by speckle, which is a specific kindof multiplicative noise. Speckle noise is a signal dependent in nature andcauses difficulties for various image interpretation tasks such as targetdetection, change detection, segmentation, edge detection, and classification. Specklenoise in SAR should be reduced while preserving important details such as edges,textures, and point scatterers in order to find the discriminative featuresneeded especially in remote sensing applications. Despeckling is generally usedas a first step in SAR image analysis and many different algorithms areproposed for that purpose. However, in the literature, generally there is atradeoff between computational load and speckle reduction quality for theproposed methods. In this paper, we present the implementation details and CPU/GPUparallelization of Sparsity Driven Despeckling (SDD) method which we proposedbefore for speckle reduction in SAR images. SDD proposes a sparsity driventotal variation (TV) approach employing l0-norm, fractional norm, orl1-norm in order to smooth homogeneous regions with minimaldegradation in edges and point scatterers. An observation image corrupted byspeckle noise is taken into consideration and SAR image despeckling problem isdefined as the optimization problem with minimization of the proposed costfunction. In order to use convex optimization methods, corresponding costfunction is approximated and written in a matrix-vector form. Thismatrix-vector form enables a special iterative optimization method where alinear system is solved in each step. Linear system which consists of positive definite symmetric matrix is solved efficiently using preconditioned conjugate gradient (PCG) iterative solver. PCGwith IC preconditioner is implemented single threaded since construction of ICpreconditioner, lower and upper Gaussian eliminations are not efficientlyparallelizable. PCG with Jacobi preconditioner is parallelized using OpenMP onCPU and CUDA on GPU since all employed operations are efficiently parallelizable.All step of the preconditioned conjugate gradient algorithm is alsoparallelized using OpenMP and CUDA. For the implementation of the PCGalgorithm, new strategies were developed to minimize the data transfer betweenCPU and GPU. We used tiling approach for large images since PCG consumes significantamount of memory and does not fit into GPU memory for large images. Therefore,in GPU parallelization streams which have an ability to perform multiple CUDAoperations simultaneously are employed. Dynamic parallelism that enables a CUDAkernel to create and synchronize new nested work is designed. We realized our GPUstudies with two versions of the developed method with and without dynamic parallelization.We performed our test studies with Nvidia GeForce and NvidiaTesla GPUs and we also used Nvidia Jetson TK1 development kit formobile platforms. In consequence, the results show that our methodprovides extremely fast (near-real-time) and high-quality SAR despeckling. Despecklingperformance, execution time and memory consumption of the proposed method areshown with detailed tables and figures using synthetic images and real-worldSAR images. Various SDD implementations for CPU and GPU are tested andinformation and suggestions are given so that reader can pick the rightimplementation strategy for a specific need.