Evaluating Logical Reasoning Ability of Large Language Models
Emunah Chan, Aldar C-F. Chan · 2026
It is observed that large language models exhibit reasoning abilities when they are sufficiently large. A key open question is whether these models carry out reasoning or simply recite memorized texts encountered during their training. We propose a new dataset comprising of unseen reasoning questions requiring semantic and deductive logical reasoning skills to evaluate the reasoning ability of large language models. The proposed evaluation framework has several desirable properties, including resilience to training data contamination, ease of answer verification, extensibility, and automated test case generation. We found that large language models exhibit reasoning ability, but this does not match the level of human intelligence. Besides, their reasoning ability is algorithmic in nature.