Towards a taxonomy of web registers and text types: a multi-dimensional analysis

Douglas Biber, Jerry Kurjian · 2007

This paper uses multi-dimensional analysis to investigate the extent to which the subject categories used by Google are linguistically well-defined. A 3.7 million word corpus is constructed by a stratified sample of web pages from two Google categories: ‘Home’ and ‘Science’. The corpus is tagged (using the Biber Tagger) and factor analysis is carried out, resulting in four factors. These factors are interpreted functionally as underlying dimensions of variation. The ‘Science’ and ‘Home’ categories are compared with respect to each dimension; although there are large differences in the dimension scores of texts within each category, the two Google categories themselves are not clearly distinguished on linguistic grounds. The dimensions are subsequently used as predictors in a cluster analysis, which identifies the ‘text types’ that are well defined linguistically. Eight text types are identified and interpreted in terms of their salient linguistic and functional characteristics.

Read the paper · More papers on PaperTik