Approches statistiques et informatiques sur l'acquisition du français langue première : une étude basée sur les suivis longitudinaux du corpus de Paris (CoLaJE)

Andrea Briglia · HAL (Le Centre pour la Communication Scientifique Directe) · 2021

CoLaJE [?] is a database composed by seven children that have been videorecorded in vivo approximately one hour every month from their first year of life until they were five. In this research, statistical treatments have been tested only on two children (Adrien and Madeleine) as for them transcription are the most complete ones. Data is transcripted in three forms: CHI is what the child says in the orthographic form, pho what the child really says and mod what he should have said according to the adult norm. To uniform data in a suitable form for automatic processing, we had to make trade-off like choices: child language is subject to interpretation difficulties by adults trying to decode it: in about 5% of the total number of occurrences, the number of words differs between the three main aforementioned forms in which sounds are coded : we decide to cut off these occurrences because they would have biased the final statistics, as the classification methods need to have an equal number of words related to the same phrase. The resulting data structure is a transformation from the video [?] into a statistically manageable database. In this respect, Code for the Human Analysis of Transcripts (CHAT) provides a standardized format for producing computerized transcripts of conversational interactions. By analyzing, cleaning, filtering and normalizing all the available original CHAT transcripts we aimed at produce two corpora composed by the overall amount of what infants said along the years, that is respectively of 8214 and 7168 annotated sentences containing more than 100 variables. Some useful measures have been calculated such as: child age in years (time); Sentence Phonetic Variation Rate (SPVR) [?]: the variation rate is obtained by comparing mod and pho in order to measure how the relation between varied and correct form evolves over time. Then, we apply a Part-Of-Speech Tagger (POS Tags), a software that reads text in a given language and assigns parts of speech to each word such as noun, verb, adjective. We used Stanford Core NLP engine [?] to tag all CHI words. 3 A brief introduction of the EM clustering method The EM clustering is an iterative method relying on the assumption that the data are generated by a mixture of underlying probability distributions, where each component represents a separate group, or cluster. The method provides the optimal number of clusters in any empirical situation, by using a two step iterative algorithm: the (E) or expectation step and the (M) or maximization step. These two steps are repeated until a further increase in the number of clusters would result in a negligible improvement in the log-likelihood, namely a convergence. Accordingly the program checks how much the overall fit improves in passing from one to two clusters (formed in all possible ways, and selecting the best), then from two to three, etc. If the error function calculated for the solution with K+1 clusters is not markedly (at least 5 percent) better than the simpler solution, with K clusters, then the solution with K clusters is considered ideal, and retained [?] [?] .To extend previous research EM Clustering method and first language acquisition 3 [?], we divide our database in strata considering 3 age classes of the child (L=1.97 - 2.64; M= 2.71 - 3.39 H=3.46 - 4.33 in years) and 3 classes of SPVR (L= 33; M=>33 and 66; H>66 in percent). So, we get 9 strata (LL, ......, HH). According to this strategy, the evolution of verbs and syntactically related forms such as articles – subject – adjectives show how morphosyntactic rules could implicitly influence the clustering procedure. We think that EM clustering method can be useful to evaluate in a reliable way linguistic structures development over time.

Read the paper · More papers on PaperTik