A Cooperative Co-learning Approach for Concept Detection in Documents

Jianqiang Li, Chunchen Liu · 2012

The learning-based approaches play a dominant role for detecting concept instances in documents, which mainly utilize the fully labeled documents (i.e., all the contained concepts are labeled) as training data to build the concept instance recognizer. However, in many cases, the available training data is sparsely labeled (only a part of the contained concepts are labeled), which makes the existing learning based approaches are not applicable. To address this issue, this paper proposed a novel co-learning approach for high accurate concept instance detection in documents. The large pool of sparsely labeled dataset is split into multiple subsets. Then multiple sequence learning models are trained on these different subsets in an iterative way, where the mechanisms of ensemble learning and co-training are embedded. The empirical experiments show that, our approach outperforms the best baselines 10% in terms of F1 measure and spend much less running time, which demonstrates the effectiveness of the proposed approach.

Read the paper · More papers on PaperTik