BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMs

Giuseppe Samo, Vivi Năstase, Chunyang Jiang, Paola Merlo · 2023

Current NLP models appear to be achieving performance comparable to human capabilities on well-established benchmarks.New benchmarks are now necessary to test deeper layers of understanding of natural languages by these models.

Read the paper · More papers on PaperTik