cleanNLP: A Tidy Data Model for Natural Language Processing

Taylor B. Arnold [aut, cre] · 2016

Provides a set of fast tools for converting a textual corpus into a set of normalized tables. Users may make use of the 'udpipe' back end with no external dependencies, or a Python back ends with 'spaCy' . Exposed annotation tasks include tokenization, part of speech tagging, named entity recognition, and dependency parsing.

Read the paper · More papers on PaperTik