Temporal Relation Classification in Hebrew
Guy Yanko, Shahaf Pariente, Kfir Bar · 2023
Temporal Relation Classification (TRC) is a fundamental task in natural language processing (NLP) and is essential for achieving a comprehensive understanding of a natural language.Given a document containing two event mentions, the objective of this task is to discern which of the two events happened first.Existing TRC datasets predominantly consist of texts written in English.To accommodate the growing interest in relevant NLP applications for Hebrew, we introduce a new TRC dataset for Hebrew.Professional annotators labeled Hebrew documents with TRC labels, adhering to guidelines adapted from a similar project on English and with some changes required to address some unique aspects of the Hebrew language.Overall, we annotated a corpus of 28,757 words, corresponding to 7,260 pairs of events.In addition to releasing the new dataset, which can be accessed at https://github.com/ shahafp/TRC-Hebrew, we train several baseline models for TRC and report their performance.