Benchmarking of semantic annotation with conditional random fields

Bruno Grilhères, C. Beauce, Stéphane Canu, Stéphan Brunessaux · 2005

The Semantic Web requires document annotation with various meta-data. But for end-users, doing it manually would be extremely time consuming and unfeasible for billion of documents. To reduce this burden, Information Extraction techniques should be applied. This paper describes the use of a recent probabilistic sequence model, Conditional Random Fields, to annotate semi-automatically sets of documents. It introduces the model principles and how to configure it to maximise the detection capabilities. The approach is evaluated on a task of event detection in news press articles related to terrorism events (the MUC-LAT corpus).

Read the paper · More papers on PaperTik