Creating A Syntactically Felicitous Constituency Treebank For Turkish

Neslihan Kara, Büşra Marşan, Merve Özçelik, Bilge Nas Arıcan, Aslı Kuzgun, Neslihan Cesur, Deniz Baran Aslan, Olcay Taner Yıldız · 2020 Innovations in Intelligent Systems and Applications Conference (ASYU) · 2020

In this study, Bakay et. al [1] and Yıldız et. al.'s [2] work on Turkish constituency treebanks were developed further. Compared to the previous work, the most prominent feature of this study is the fact that every annotation and refinement process is held manually. In addition, constituency treebank created as a result of this study abides by the syntactic rules and typologic features of Turkish while the trees created by previous studies convey only the translated and simply inverted trees that completely ignore the syntactic properties of Turkish. The methodology followed in this study resulted in a significantly more accurate representation of Turkish language and simpler, relatively flatter trees. The straightforward style of trees in this study reduces the complexity and offers a better training dataset for learning algorithms.

Read the paper · More papers on PaperTik