Nonsymbolic Text Representation
Hinrich Schütze · 2017
We introduce the first generic text representation model that is completely nonsymbolic, i.e., it does not require the availability of a segmentation or tokenization method that attempts to identify words or other symbolic units in text.This applies to training the representations as well as to using them in an application.We demonstrate better performance than prior work on entity typing and text denoising.