Relating linguistic units to socio-contextual information in a spontaneous speech corpus of Spanish
José María Guirao, Antonio Moreno Sandoval, Ana González Ledesma, Guillermo de la Madrid, Manuel Alcántara · 2006
This chapter shows the application of statistical tests to a corpus of spontaneous spoken Spanish. Our goal is to find representative differences between different parts of the corpus. To this end, we tagged n-grams in the corpus with features related to the speaker (age, gender, etc.), or the context (dialogue, monologue, media, etc.), and applied the log-likelihood test (Dunning, 1993) in order to find the most distinctive lexical or grammatical items for each specific socio-contextual feature. This chapter is divided in three sections. In the first, the characteristics of the spoken corpus are shown. The second section is devoted to the explanation of the computational tool. In the third section, a first rough estimate of the results obtained is given, as well as possible applications of the model.