A Practical Big Data Use Case

Églantine Schmitt · Big Data · 2020

This chapter describes how the problems of corpus constitution and data cleaning concretely materialize in a given situation. It presents a case study in an organization, the software publisher Proxem, that processes massive amounts of data. In terms of mythology as in terms of practice, Proxem is a player in the computational processing of digital footprints. The epistemological gesture of data aggregation and commensuration is coupled with a set of technical gestures by which textual data are effectively made manipulable by Proxem employees. The constitution of the coding plan is therefore an art which mobilizes a set of linguistic but also extra linguistic skills. The effective classification of feedback in the coding plan is not done unitarily, but through the composition of linguistic rules based on the presence or absence of certain words, their order, their nature, etc. The quantification of codes is a way to identify regularities, a statistical standard.

Read the paper · More papers on PaperTik