An analysis of the joint venture Japanese text prototype and its effect on system performance

Steve Maiorano · 1993

The TIPSTER Data Extraction and Fifth Message Understanding Conference (MUC-5) tasks focused on the process of data extraction. This is a procedure in which prespecified types of information are identified within free text, extracted, and inserted automatically within a template. Three TIPSTER contractors -- BBN, GE/CMU, NMSU/Brandeis -- participated in the August '93 MUC-5 evaluation for both the English joint venture (EJV) and English microelectronics (EME) domains and their Japanese-language counterparts, the JJV and JME applications. Two other contractors -- SRI and SRA -- participated in the EJV and JJV domains alone. CMU's Textract system took part in the Japanese-language domains only. Of the five systems that tested in both English and Japanese, all but one scored higher in the Japanese-language applications according to both the summary error-based scores and recall/precision-based metrics. This overall result has lead some participants and observers to suggest that Japanese is an "easier" language than English.

Read the paper · More papers on PaperTik