Evaluating Transformer’s Ability to Learn Mildly Context-Sensitive Languages

Shunjie Wang, Shane Steinert‐Threlkeld · 2023

Despite the fact that Transformers perform well in NLP tasks, recent studies suggest that selfattention is theoretically limited in learning even some regular and context-free languages.These findings motivated us to think about their implications in modeling natural language, which is hypothesized to be mildly contextsensitive.We test the Transformer's ability to learn mildly context-sensitive languages of varying complexities, and find that they generalize well to unseen in-distribution data, but their ability to extrapolate to longer strings is worse than that of LSTMs.Our analyses show that the learned self-attention patterns and representations modeled dependency relations and demonstrated counting behavior, which may have helped the models solve the languages.

Read the paper · More papers on PaperTik