Mobile phone identification using recorded speech signals
Constantine L. Kotropoulos, Stamatios Samaras · 2014
In this paper, we elaborate on mobile phone identification from recorded speech signals. The goal is to extract intrinsic traces related to the mobile phone used to record a speech signal. Mel frequency cepstral coefficients (MFCCs) are extracted from any recorded speech signal at a frame level. The sequences of the MFCC vectors extracted from each recording device train a Gaussian Mixture Model with diagonal covariance matrices. A Gaussian supervector is derived by concatenating the mean vectors and the main diagonals of the covariance matrices that is used as a template for each device. Experiments were conducted on a database of 21 mobile phones of various models from 7 different brands. The aforementioned database, that is called MOBIPHONE, was collected by recording 10 utterances, uttered by 12 male speakers and another 12 female speakers, randomly chosen from the TIMIT database. Three commonly used classifiers were employed, such as Support Vector Machines with different kernels, a Radial Basis Functions neural network, and a Multi-Layer Perceptron. The best identification accuracy (97.6%) was obtained by the Radial Basis Functions neural network.