A roadmap towards determining the universal status of semantic frames

Hans Christian Boas · 2020

The Berkeley FrameNet project, founded in 1997, organizes the lexicon of English by semantic frames (Fillmore 1982), with valence information derived from attested, manually annotated corpus examples. The resulting FrameNet database contains more than one thousand frames, together with more than twelve thousand lexical unites and close to 200,000 annotated example sentences. FrameNet data have been used to answer a variety of empirical research questions on the mapping from semantics to syntax and they have been employed in a number of NLP tasks such as role labeling and text summarization. Since the early 2000s, several projects have re-used the semantic frames based on English for constructing FrameNets for other languages, most notably Spanish, Japanese, German, and Swedish, among others. While the tools, corpora, and databases differ from each other, the main organizing principle, the semantic frame, used for structuring the lexicon remains similar across all the FrameNets for different languages. The motivation for re-using semantic frames from English for other languages is the idea that frames are universal, similar to Fillmore’s original case roles. However, there has not yet been any empirical investigation into what constitutes “universal” frames or how one can possibly determine the universal status of semantic frames. This paper proposes a systematic method for identifying semantic frames that could be labeled “universal” (based only on data from languages under investigation). We specifically address the question of how semantic frames can be used for contrastive analysis.

Read the paper · More papers on PaperTik