High throughput VLSI architecture for gradient guided filter with approximated arithmetic operations

Lei Wu · 2018

Guided image filtering has been applied widely for increasing demand of high performance filtering, especially for real-time image/video processing.Gradient guided filter improves the filtering quality, reducing the halo-artifacts problem due to its edge-aware characteristics.However, the gradient guided filter algorithm has high computation complexity and the computation involves global pixels, which hinder its VLSI implementation for real-time full-HD application.This work addresses these issues and a VLSI architecture is proposed for the gradient guided image filter.Several design techniques are developed and used in the design to achieve high computation speed and high throughput.A seamless dataflow is proposed for the complete system that consists of three main processing stages, specifically the preprocessing stage, the linear coefficient computation stage and the output stage.The preprocessing stage applies a down-sampling technique with a large sampling rate to reduce computation cost in terms of circuit size, processing time and power consumption.The global parameter values are quickly derived with reasonable good global information maintained so that the quality of the filtering results are not sacrificed when these values are applied in the subsequent two stages.The linear coefficient computation stage contains the most complex computations such as square root, division and exponential function.Down-sampling technique is applied with a sampling rate lower than in the preprocessing stage so as to balance the computation cost and filtering accuracy.In addition, the intensive arithmetic computation modules that dominate the critical path delay are redesigned by using adequate approximated operations.Specifically, novel non-iterative dividers are developed to replace original dividers for reducing delays in the critical paths.With the proposed non-iterative division, MSB Most Significant Bit MAC multiplier-accumulator MAEP Maximum Absolute Error Percentage SQRT SQuare RooT VLSI Very Large Scale Integration WLS Weighted Least Squares xiii

Read the paper · More papers on PaperTik