Stacked Autoencoder Based HRTF Synthesis from Sparse Data

Sunil Bharitkar, Timothy A. Mauer, Teresa Wells, David M. Berfanger · 2018

Ipsilateral and contralateral head-related transfer functions (HRTF),$H_{\mathbf{ipsi}}(\omega, r,\phi,\theta)$and$H_{\mathbf{contra}}(\omega, r,\phi,\theta)$, are used for creating the perception of a virtual sound source at an arbitrary distance$r$and azimuth-elevation tuple$\underline{\psi}=[\theta,\phi]^{T}$relative to the median plane for a given frequency$\omega$. Publicly available databases use a subset of a full-grid of angular directions due to time and complexity to acquire and deconvolve responses. In this paper, we present a subspace-based technique for reconstructing HRTFs at arbitrary directions for the IRCAM-Listen HRTF database, which comprises a sparse set of HRTFs sampled every 15° along the azimuth/elevation direction. The presented technique includes first augmenting the sparse IRCAM dataset using auditory localization blur, then deriving a set of lower-dimensional compressed representation (using an autoencoder) from the augmented HRTFs. The lower dimensional representations are then trained using a fully-connected neural network (FCNN) for the corresponding directions. The reconstruction of HRTF corresponding to an arbitrary direction$\underline{\psi}_{p}$is achieved by applying the compressed output from the FCNN, for an arbitrary direction, to a reconstruction system (viz., a decoder of an autoencoder). The results demonstrate the autoencoder approach provides good quality objective and subjective results.

Read the paper · More papers on PaperTik