BLM-s/lE: A structured dataset of English spray-load verb alternations for testing generalization in LLMs
Giuseppe Samo, Vivi Năstase, Chunyang Jiang, Paola Merlo · 2023
Current NLP models appear to be achieving performance comparable to human capabilities on well-established benchmarks.New benchmarks are now necessary to test deeper layers of understanding of natural languages by these models.