A Method to Lexical Normalisation of Tweets
Pablo Gamallo, Marcos García, José Ramom Pichel Campos · 2013
espanolEste articulo describe una estrategia de normalizacion lexica de palabras “out-of-vocabulary” (OOV) en tweets escritos en espanol. Para corregir OOV incorrectos, el sistema de normalizacion genera candidatos “in-vocabulary” (IV) que aparecen en diferentes recursos lexicos y selecciona el mas adecuado. Nuestro metodo genera dos tipos de candidatos, primarios y secundarios, que seran ordenados de diferentes maneras en el proceso de seleccion del mejor candidato. EnglishThis paper describes a strategy to perform lexical normalisation of out- of-vocabulary (OOV) words in Spanish tweets. To correct any ill-formed OOV, the normalisation system generates in-vocabulary (IV) candidates found in several lexical resources, and selects the best one. Our method generates two types of candidates, primary and secondary IV candidates, which will be ranked in different ways to select the best candidate.