LLM Reasoning: from OpenAI O1 to DeepSeek R1

Jiaqi Wang, Xinliang Li, Zhengliang Liu, Wu Zihao, Zhong Tianyang, Peng Shu, Yiwei, Li, Jiang Hanqi, Yifan, Zhou, Junhao Chen, Ruan Wei, Yi Feng Pan, Zhao Huaqin, Chong Ma, Zhenyuan, Yang, Xu Shaochen, Zhang Ruidong, Dai Haixing, Lin Zhao, Dehong Gao · HAL (Le Centre pour la Communication Scientifique Directe) · 2025

This review comprehensively explores reasoning capabilities in large language models (LLMs), tracing the evolution from OpenAI's O1 to DeepSeek's R1 while analyzing the underlying technological innovations. We examine three fundamental mechanisms enabling LLM reasoning: ICL through prompt engineering, architectural innovations including attention mechanisms and sparse architectures, and reinforcement learning (RL) strategies such as Group Relative Policy Optimization (GRPO). The paper compares OpenAI's reinforcement learning from human feedback-based approach with DeepSeek's direct RL methodology, highlighting how the latter achieves comparable performance with greater computational efficiency. Futheremore, our analysis extends to model compression techniques including quantization and pruning, knowledge distillation (KD) strategies that democratize access to reasoning capabilities, and system-level optimizations for inference. Additionally, we discuss the integration of these models into AI agents, ethical considerations surrounding their deployment, and current limitations including reasoning-induced hallucinations and inefficient resource allocation between simple and complex tasks. By synthesizing theoretical foundations with practical applications across domains like healthcare and education, this paper provides valuable guidance for researchers and practitioners working to enhance reasoning capabilities in LLMs while ensuring their responsible deployment.

Read the paper · More papers on PaperTik