Fine-Grained Reasoning Evaluation: How Well Does DeepSeek-R1 Handle Causal Inference?
Kamal Singh Bisht · Technix International Journal for Engineering Research · 2025
The evaluation of DeepSeek-R1, a reinforcement learning-fine-tuned open-source model, on causal inference tasks reveals distinct patterns of strength and limitation across different dimensions of causal reasoning. Using the novel CausalQA-R benchmark spanning healthcare, finance, policy, and commonsense domains, the investigation demonstrates that while DeepSeek-R1 excels at direct causal identification and maintains strong causal consistency, it struggles with multi-hop causal chains, confounding variables, and counterfactual scenarios. Performance notably degrades when moving up Pearl's causal hierarchy from association to intervention to counterfactuals. The model shows particular strength in handling probabilistic causation but exhibits overreliance on explicit causal language rather than deeper structural understanding. These observations contribute to theoretical discussions about whether language models can develop true causal understanding through scale alone or require fundamental architectural innovations tailored specifically to causal cognition.