Feature-Rich Information Extraction for the Technical Trend-Map Creation.
Risa Nishiyama, Yuta Tsuboi, Yuya Unno, Hironori Takeuchi · 2010
The authors used a word sequence labeling method for technical effects and base-technology extraction in the Technical Trend Map Creation Subtask of the NTCIR-8 Patent Mining Task. The method labels each word based on CRF (Conditional Random Field) trained with labeled data. The word features employed in the labeling are obtained by using explicit/implicit document structures, technology fields assigned to the document, effect context phrases, phrase de-pendency structures and a domain adaptation technique. Results of the formal run showed that the explicit document structure feature and the phrase dependency structure feature are effective in anno-tating patent data. The implicit document structure feature and the domain adaptation feature are also effective for annotating paper data.