Cocytus : parallel Natural Language Processing
Noah Evans · Institutional Repositories DataBase (IRDB) · 2008
Natural Language Processing(NLP) is at a crossroads.As linguistic data sets increase in size the ability of NLP tools to run quickly enough to process these data sets is not growing fast enough to compensate.This inability of tool speed to keep up to data size is compounded by the computationally intensive nature of modern statistical NLP algorithms.This need to handle large data systems makes systems that can solve NLP problems over large datasets in a scalable way paramount.There are currently tools trying to address this scalability problem, but they can't address scalability problems in either a cost effective way or a way that is tenable given the resources of a typical researcher.This thesis solves the scalability problem by developing a Software Architecture for Language Engineering(SALE) that simplifies solving NLP problems over large data sets by transparently representing computation and various linguistic resources in a way that is transparent to the user.It does this by developing two systems, a set of data transformers that present data in a variety of NLP structured data formats and character sets in new format called TreePaths and utf8 respectively, and an implementation of the MapReduce parallel algorithm which automatically divides linguistic resources and spreads them over a variety of computational resources using heterogeneous operating systems.This allows the system to deal with data in many more formats, in larger sizes more easily than current methods.