A Derivation of the Soft-Thresholding Function
Ivan Selesnick · 2013
These notes show the derivation of non-linear soft-thresholding function for signal denoising. The softthresholding function can be used for denoising by applying it to the transform-domain representation, provided the transform yields a sparse representation of the signal. For example, wavelet transforms provide sparse representations of piece-wise smooth signals, and the short-time Fourier transform (STFT) provides sparse representations of oscillatory signals (like speech). 1 Derivation of the Soft-Threshold Function We assume that a signal of interest has been corrupted by additive noise, i.e. g = x + n (1) where n is white zero-mean Gaussian noise independent of the signal x. We observe g (a noisy signal), and wish to estimate the noise-free signal x as accurately as possible. In the transform domain (eg, wavelet, STFT, etc), the problem can be formulated as y = w + n (2) where y is the noisy coefficient (eg, wavelet coefficient), w is the noise-free coefficient and n is noise, which is again zero-mean Gaussian (any linear transform of a zero-mean Gaussian random signal results in a zeromean Gaussian random signal). If the transform is orthogonal, then the noise in the transform domain has the same correlation function as the original noise in the signal domain; therefore, when the transform is orthogonal, white noise in the signal domain becomes white noise in the transform domain. Our goal is to estimate w from the noisy observation y. The estimate will be denoted as ˆw. Because the estimate depends on the observed (noisy) value y, we also denote the estimate as ˆw(y). We will use the maximum a posteriori (MAP) estimator. The MAP estimator is based on the probability density function (pdf) of w. Specifically, given an observed value y, the MAP estimator asks what value of w is most likely? That is, the MAP estimator looks for the value of w where the probability of w is highest; it looks for the peak value. Therefore, the MAP estimator is defined as 1 ˆw(y) = arg max w p w|y(w|y) (3) where ‘arg max ’ is the value of the argument where the function has its maximum. The pdf p w|y(w|y) is the distribution of w given a specific value y. Conditional pdf, p w|y(w|y), for some value of y. and −10 −8 −6 −4 −2 0 2 4 6 8 10