A statistical profile of the Named Entity task
David D. Palmer, David S. Day · 1997
In this paper we present a statistical profile of the Named Entity task, a specific information extraction task for which corpora in several languages are available.Using the results of the statistical analysis, we propose an algorithm for lower bound estimation for Named Entity corpora and discuss the significance of the cross-lingual comparisons provided by the analysis.Portuguese corpora 1 which were prepared according to the language, as well as a breakdown of total phrases into the MET guidelines.Table 1 shows the sources for the corpora, three individual categories.