Unicode and Natural Language

Moritz Lenz · Apress eBooks · 2017

Text that computers deal with tends to fall into two categories: things that are meant to be consumed by humans (like prose), and things that are meant to be consumed by software (machine code and encrypted files come to mind). These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik