The UAM CorpusTool: software for corpus annotation and exploration

Michael O'Donell · 2009

This paper describes the capabilities of the UAM CorpusTool, software for the annotation of text corpora. The software allows the user to annotate a corpus of text files at a number of linguistic layers, which are defined by the user. For instance, one can annotate texts at the document layer (e.g., text type, writer characteristics, register, etc.), semantic-pragmatic levels, and at syntactic levels (e.g., clause, phrases, etc.). At each annotation layer, the user defines a hierarchy of tags appropriate for that layer, using a graphical tool. The user then annotates the text at each layer by swiping the text to indicate a segment, and then assigning features from the tag hierarchy at that layer. The paper also descibes some of the supporting functionalites of the software, including corpus search, automatic tagging based on lexical pattern matching, and production of statistics. RESUMEN : Este articulo describe las caracteristicas del UAM CorpusTool, un software para la anotacion de corpus de texto. Este software permite al usuario anotar un corpus de archivos de texto en distintos niveles linguisticos previamente definidos por el usuario. Por ejemplo, uno puede anotar los textos a nivel de documento (Ej., tipo de texto, caracteristicas del escritor, registro, etc.), en niveles semantico-pragmaticos, y en niveles sintacticos (Ej., clausula, sintagmas, etc.). Utilizando una herramienta grafica, el usuario define una jerarquia de etiquetas apropiadas para cada nivel de anotacion. A continuacion, el usuario anota el texto en cada nivel, seleccionando primero el texto para indicar un segmento, y asignandole caracteristicas elegidas de entre la jerarquia de etiquetas definidas para ese nivel. Este articulo describe tambien otras funcionalidades anadidas al software, tales como instrumento de busqueda en el corpus, etiquetador automatico basado en la correspondencia de patrones lexicos, y produccion de informes estadisticos.

Read the paper · More papers on PaperTik